Wei Zeng 0004

dblp:80/1961-4 · DBLP profile ↗
← Back
76ranked-venue papers
10as first author
61since 2021 · last 2026
0000-0002-5600-8824ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 7 first-author · 32 since 2021Human-computer interaction and ubiquitous computing · 27 · 1 first-author · 23 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 Do Large Language Models Reason About Uncertainty Like Humans? A Benchmark on Hurricane Forecast Visualization Comprehension
abstract
Uncertainty visualizations, such as hurricane cones and ensemble tracks, are essential for risk communication but are often misinterpreted, leading to harmful decisions. As AI assistants like large language models (LLMs) increasingly support understanding of graphics and decision-making, they offer a promising pathway to enhance the interpretation of complex visualizations and a new opportunity to examine and improve the interpretation of uncertainty. We introduce UnReason, the first benchmark that systematically compares how humans and LLMs reason about hurricane forecast uncertainty visualizations. UnReason spans two escalating phases, seven representative visualization formats, six real hurricane cases, and three agent types (humans, LLMs with context, and LLMs without context), including 880 visualizations and 117,600 structured question–answer pairs under matched evaluation conditions. Phase 1 evaluates reasoning across implicit and explicit uncertainty encodings; Phase 2 examines reasoning under single- versus multi-dimensional uncertainty representations. We thoroughly assess damage estimation, reasoning strategies, and comprehension patterns, revealing that LLMs have a stronger semantic and conceptual understanding of uncertainty, and are less misled by visual variability, but still replicate key human biases during decision-making. Our findings offer insights into aligning LLM behavior with human cognition in uncertainty-rich visual reasoning tasks.
Le Liu 0008, Bohan Shen, Wei Zeng 0004, Shizhou Zhang, Di Xu 0010, Peng Wang 0015
AAAI4
2026 Agency-Enhanced Visual Search in VR: Robust to Distraction, Delay, and Perspective Shifts
abstract
The sense of agency, the internal feeling of controlling one's actions and their outcomes, fundamentally shapes user interaction in virtual environments. Although extensively studied for its subjective impact, agency's implicit yet critical role in guiding visual-spatial attention has been largely overlooked. This study investigated whether the sense of agency, conferred by prior active control, enhances subsequent visual search efficiency for the previously controlled stimulus under conditions of attentional, temporal, and spatial perturbation. Our results show that agency-driven attentional benefits are remarkably robust, persisting despite competing salient visual distractors, delayed outcomes, and changes in spatial layout. Furthermore, when a delay was introduced between the control and visual search, the agency effect attenuated for targets presented beyond the operators’ peripersonal space. These findings provide valuable insights for constructing immersive user experiences and advancing theoretical frameworks in human-computer interaction, suggesting efficient strategies to support sustained user engagement in virtual and augmented reality.
Chunlin Liao, Wei Zeng 0004, Rongrong Chen 0004
CHI2
2026 InkIdeator: Supporting Chinese-Style Visual Design Ideation via AI-Infused Exploration of Chinese Paintings
abstract
Visual designers often seek inspiration from Chinese paintings when tasked with creating Chinese-style illustrations, posters, etc. Our formative study (N=10) reveals that during ideation, designers learn the cultural symbols, emotions, compositions, and styles in Chinese paintings but face challenges in searching, analyzing, and integrating these dimensions. This paper leverages multi-modal large models to annotate the value of each dimension in 16,315 Chinese paintings, built on which we propose InkIdeator, an ideation support system for Chinese-style visual designs. InkIdeator suggests cultural symbols associated with the task theme, provides dimensional keywords to help analyze Chinese paintings, and generates visual examples integrating user-selected keywords. Our within-subjects study (N=12) using a baseline system without extracted dimensional keywords, along with two extended use cases by Chinese painters, indicates InkIdeator's effectiveness in creative ideation support, helping users efficiently explore cultural dimensions in Chinese paintings and visualize their ideas. We discuss implications for supporting culture-related visual design ideation with generative AI.
Ziyao Gao, Zhendong He, Zongtan He, Zhupeng Huang, Wei Zeng 0004, Xiaojuan Ma, Zhenhui Peng
CHI7
2026 Introducing ManyViews: an AI-assisted tool to support citizens' engagement in the design of urban spaces
Rong Huang 0007, Yihan Hou, Mela Bettega, Kang Zhang 0001, Wei Zeng 0004
Int. J. Hum. Comput. Stud.5
2026 DKMap: Interactive Exploration of Vision-Language Alignment in Multimodal Embeddings via Dynamic Kernel Enhanced Projection
abstract
Examining vision-language alignment in multimodal embeddings is crucial for various tasks, such as evaluating generative models and filtering pretraining data. The intricate nature of high-dimensional features necessitates dimensionality reduction (DR) methods to explore alignment of multimodal embeddings. However, existing DR methods fail to account for cross-modal alignment metrics, resulting in severe occlusion of points with divergent metrics clustered together, inaccurate contour maps from over-aggregation, and insufficient support for multi-scale exploration. To address these problems, this paper introduces DKMap, a novel DR visualization technique for interactive exploration of multimodal embeddings through Dynamic Kernel enhanced projection. First, rather than performing dimensionality reduction and contour estimation sequentially, we introduce a kernel regression supervised t-SNE that directly integrates post-projection contour mapping into the projection learning process, ensuring cross-modal alignment mapping accuracy. Second, to enable multi-scale exploration with dynamic zooming and progressively enhanced local detail, we integrate validation-constrained a refinement of a generalized t-kernel with quad-tree-based multi-resolution technique, ensuring reliable kernel parameter tuning without overfitting. DKMap is implemented as a multi-platform visualization tool, featuring a web-based system for interactive exploration and a Python package for computational notebook analysis. Quantitative comparisons with baseline DR techniques demonstrate DKMap's superiority in accurately mapping cross-modal alignment metrics. We further demonstrate generalizability and scalability of DKMap with three usage scenarios, including visualizing million-scale text-to-image corpus, comparatively evaluating generative models, and exploring a billion-scale pretraining dataset.
Chenxi Ruan, Yu Zhang 0043, Zikun Deng, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.5
2025 Poll-Sketcher: Visual Exploration of Time-Varying Air Pollutant Data Based on Hand-Drawn Sketches
Jianing Hao, Jingxuan Feng, Chongke Bi, Xiaobin Qiu, Wei Zeng 0004
CGI (3)7
2025 GenColor: Generative Color-Concept Association in Visual Design
abstract
Existing approaches for color-concept association typically rely on query-based image referencing, and color extraction from image references. However, these approaches are effective only for common concepts, and are vulnerable to unstable image referencing and varying image conditions. Our formative study with designers underscores the need for primary-accent color compositions and context-dependent colors (e.g., 'clear' vs. 'polluted' sky) in design. In response, we introduce a generative approach for mining semantically resonant colors leveraging images generated by text-to-image models. Our insight is that contemporary text-to-image models can resemble visual patterns from large-scale real-world data. The framework comprises three stages: concept instancing produces generative samples using diffusion models, text-guided image segmentation identifies concept-relevant regions within the image, and color association extracts primarily accompanied by accent colors. Quantitative comparisons with expert designs validate our approach's effectiveness, and we demonstrate the applicability through cases in various design scenarios and a gallery.
Yihan Hou, Xingchen Zeng, Yusong Wang 0004, Manling Yang, Wei Zeng 0004
CHI6
2025 SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches
Haichuan Lin, Jiazhi Xia, Wei Zeng 0004
CHI4
2025 AKRMap: Adaptive Kernel Regression for Trustworthy Visualization of Cross-Modal Embeddings
abstract
Cross-modal embeddings form the foundation for multi-modal models. However, visualization methods for interpreting cross-modal embeddings have been primarily confined to traditional dimensionality reduction (DR) techniques like PCA and t-SNE. These DR methods primarily focus on feature distributions within a single modality, whilst failing to incorporate metrics (e.g., CLIPScore) across multiple modalities. This paper introduces AKRMap, a new DR technique designed to visualize cross-modal embeddings metric with enhanced accuracy by learning kernel regression of the metric landscape in the projection space. Specifically, AKRMap constructs a supervised projection network guided by a post-projection kernel regression loss, and employs adaptive generalized kernels that can be jointly optimized with the projection. This approach enables AKRMap to efficiently generate visualizations that capture complex metric distributions, while also supporting interactive features such as zoom and overlay for deeper exploration. Quantitative experiments demonstrate that AKRMap outperforms existing DR methods in generating more accurate and trustworthy visualizations. We further showcase the effectiveness of AKRMap in visualizing and comparing cross-modal embeddings for text-to-image models. Code and demo are available at https://github.com/yilinye/AKRMap.
Junchao Huang, Xingchen Zeng, Jiazhi Xia, Wei Zeng 0004
ICML5
2025 nvBench 2.0: Resolving Ambiguity in Text-to-Visualization through Stepwise Reasoning
abstract
Text-to-Visualization (Text2VIS) enables users to create visualizations from natural language queries, making data insights more accessible. However, Text2VIS faces challenges in interpreting ambiguous queries, as users often express their visualization needs in imprecise language. To address this challenge, we introduce nBench 2.0, a new benchmark designed to evaluate Text2VIS systems in scenarios involving ambiguous queries. nvBench 2.0 includes 7,878 natural language queries and 24,076 corresponding visualizations, derived from 780 tables across 153 domains. It is built using a controlled ambiguity-injection pipeline that generates ambiguous queries through a reverse-generation workflow. By starting with unambiguous seed visualizations and selectively injecting ambiguities, the pipeline yields multiple valid interpretations for each query, with each ambiguous query traceable to its corresponding visualization through step-wise reasoning paths.We evaluate various Large Language Models (LLMs) on their ability to perform ambiguous Text2VIS tasks using nBench 2.0. We also propose Step-Text2Vis, an LLM-based model trained on nvBench 2.0, which enhances performance in ambiguous scenarios through step-wise preference optimization. Our results show that Step-Text2Vis outperforms all baselines, setting a new state-of-the-art for ambiguous Text2VIS tasks. Our source code and data are available at https://nvbench2.github.io/
Tianqi Luo, Chuhan Huang, Leixian Shen, Boyan Li 0001, Shuyu Shen, Wei Zeng 0004, Nan Tang 0001, Yuyu Luo
NeurIPS6
2025 HeritageExplorer: Interactive Visualization and Dialogue System for Multi-modal Architectural Heritage Exploration
abstract
Effective visualization is essential for cultural heritage interpretation. However, existing visualization systems remain constrained by fragmented data integration and limited exploration capabilities for multimodal heritage data. This paper presents HeritageExplorer, an interactive system that synergizes large language models (LLMs) with dynamic visualizations to enable progressive heritage exploration. Our approach constructs a comprehensive knowledge graph integrating 831 historic buildings in Guangzhou, which unifies their architectural, spatial, temporal, and contextual attributes. The system’s novel integration of KG-enhanced contextual understanding with LLMs supports: natural language query understanding and seamless coupling with interactive visualizations. Quantitative evaluation demonstrates consistent improvements in factual accuracy across heritage tasks, while case studies illustrate its successful application in diverse exploration scenarios.
Yusong Wang 0004, Yihan Hou, Rong Huang 0007, Wei Zeng 0004
VINCI4
2025 SceneWeaver: A Multi-Agent Collaborative System for 3D Scene Creation in Video Games
abstract
Scene creation in video games is an iterative process that transitions from initial ideas to 2D visuals and finally to 3D scenes, requiring collaborative teamwork among gameplay and environment designers, 3D modelers, and technical artists. While advances in generative methods have enabled the automation of scene creation, these end-to-end approaches often do not align with the typical workflow. This paper presents SceneWeaver, a 3D scene creation system that utilizes a multi-agent collaborative framework with large language models (LLMs) assigned to manage text parsing, floorplan design, object selection, and scene composition. An interactive visual interface with step-by-step previsualization to enhance human-AI communication and collaboration throughout the scene creation process. SceneWeaver effectively enhances user engagement and improves process controllability in 3D scene creation, as confirmed by performance evaluations and user studies.
Rong Huang 0007, Chenxi Ruan, Bingchuan Jiang, Wei Zeng 0004
VINCI4
2025 VizTA: Enhancing Comprehension of Distributional Visualization with Visual-Lexical Fused Conversational Interface
abstract
Abstract Comprehending visualizations requires readers to interpret visual encoding and the underlying meanings actively. This poses challenges for visualization novices, particularly when interpreting distributional visualizations that depict statistical uncertainty. Advancements in LLM‐based conversational interfaces show promise in promoting visualization comprehension. However, they fail to provide contextual explanations at fine‐grained granularity, and chart readers are still required to mentally bridge visual information and textual explanations during conversations. Our formative study highlights the expectations for both lexical and visual feedback, as well as the importance of explicitly linking these two modalities throughout the conversation. The findings motivate the design of VizTA, a visualization teaching assistant that leverages the fusion of visual and lexical feedback to help readers better comprehend visualization. VizTA features a semantic‐aware conversational agent capable of explaining contextual information within visualizations and employs a visual‐lexical fusion design to facilitate chart‐centered conversation. A between‐subject study with 24 participants demonstrates the effectiveness of VizTA in supporting the understanding and reasoning tasks of distributional visualization across multiple scenarios.
Liangwei Wang 0001, Zhan Wang 0001, Shishi Xiao, Le Liu 0008, Fugee Tsung, Wei Zeng 0004
Comput. Graph. Forum6
2025 NFTracer: Tracing NFT Impact Dynamics in Transaction-Flow Substitutive Systems With Visual Analytics
abstract
Impact dynamics are crucial for estimating the growth patterns of NFT projects by tracking the diffusion and decay of their relative appeal among stakeholders. Machine learning methods for impact dynamics analysis are incomprehensible and rigid in terms of their interpretability and transparency, whilst stakeholders require interactive tools for informed decision-making. Nevertheless, developing such a tool is challenging due to the substantial, heterogeneous NFT transaction data and the requirements for flexible, customized interactions. To this end, we integrate intuitive visualizations to unveil the impact dynamics of NFT projects. We first conduct a formative study and summarize analysis criteria, including substitution mechanisms, impact attributes, and design requirements from stakeholders. Next, we propose the Minimal Substitution Model to simulate substitutive systems of NFT projects that can be feasibly represented as node-link graphs. Particularly, we utilize attribute-aware techniques to embed the project status and stakeholder behaviors in the layout design. Accordingly, we develop a multi-view visual analytics system, namely NFTracer, allowing interactive analysis of impact dynamics in NFT transactions. We demonstrate the informativeness, effectiveness, and usability of NFTracer by performing two case studies with domain experts and one user study with stakeholders. The studies suggest that NFT projects featuring a higher degree of similarity are more likely to substitute each other. The impact of NFT projects within substitutive systems is contingent upon the degree of stakeholders' influx and projects' freshness.
Yifan Cao 0001, Lue Shen, Kani Chen, Yang Wang 0020, Wei Zeng 0004, Huamin Qu
IEEE Trans. Vis. Comput. Graph.6
2025 FinFlier: Automating Graphical Overlays for Financial Visualizations With Knowledge-Grounding Large Language Model
abstract
Graphical overlays that layer visual elements onto charts, are effective to convey insights and context in financial narrative visualizations. However, automating graphical overlays is challenging due to complex narrative structures and limited understanding of effective overlays. To address the challenge, we first summarize the commonly used graphical overlays and narrative structures, and the proper correspondence between them in financial narrative visualizations, elected by a survey of 1752 layered charts with corresponding narratives. We then design FinFlier, a two-stage innovative system leveraging a knowledge-grounding large language model to automate graphical overlays for financial visualizations. The text-data binding module enhances the connection between financial vocabulary and tabular data through advanced prompt engineering, and the graphics overlaying module generates effective overlays with narrative sequencing. We demonstrate the feasibility and expressiveness of FinFlier through a gallery of graphical overlays covering diverse financial narrative visualizations. Performance evaluations and user studies further confirm system's effectiveness and the quality of generated layered charts.
Jianing Hao, Manling Yang, Yuzhe Jiang, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.6
2025 Dashboard Vision: Using Eye-Tracking to Understand and Predict Dashboard Viewing Behaviors
abstract
Dashboards serve as effective visualization tools for conveying complex information. However, there exists a knowledge gap regarding how dashboard designs impact user engagement, necessitating designers to rely on their design expertise. Saliency has been used to comprehend viewing behaviors and assess visualizations, yet existing saliency models are primarily designed for single-view visualizations. To address this, we conduct an eye-tracking study to quantify participants' viewing patterns on dashboards. We collect eye-movement data from 60 participants, each viewing 36 dashboards (16 representative dashboards shared across all and 20 unique to each participant), totaling 1,216 dashboards and 2,160 eye-movement data instances. Analysis of the data from 16 dashboards viewed by all participants provides insights into how dashboard objects and layout designs influence viewing behaviors. Our analysis confirms known viewing patterns and reveals new patterns related to dashboard layout designs. Using the eye-movement data and identified patterns, we develop a saliency model to predict viewing behaviors with dashboards. Compared to state-of-the-art models for single-view visualizations, our model demonstrates overall improvement in prediction performance for dashboards. Finally, we propose potential dashboard design guidelines, illustrate an application case, and discuss general scanning strategies along with limitations and future work.
Manling Yang, Yihan Hou, Remco Chang, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.5
2025 ModalChorus: Visual Probing and Alignment of Multi-Modal Embeddings via Modal Fusion Map
abstract
Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting in decreased model performance and diminished generalization. To address this problem, we design ModalChorus, an interactive system for visual probing and alignment of multi-modal embeddings. ModalChorus primarily offers a two-stage process: 1) embedding probing with Modal Fusion Map (MFM), a novel parametric dimensionality reduction method that integrates both metric and nonmetric objectives to enhance modality fusion; and 2) embedding alignment that allows users to interactively articulate intentions for both point-set and set-set alignments. Quantitative and qualitative comparisons for CLIP embeddings with existing dimensionality reduction (e.g., t-SNE and MDS) and data fusion (e.g., data context map) methods demonstrate the advantages of MFM in showcasing cross-modal features over common vision-language datasets. Case studies reveal that ModalChorus can facilitate intuitive discovery of misalignment and efficient re-alignment in scenarios ranging from zero-shot classification to cross-modal retrieval and generation.
Shishi Xiao, Xingchen Zeng, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.4
2025 Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
abstract
Emerging multimodal large language models (MLLMs) exhibit great potential for chart question answering (CQA). Recent efforts primarily focus on scaling up training datasets (i.e., charts, data tables, and question-answer (QA) pairs) through data collection and synthesis. However, our empirical study on existing MLLMs and CQA datasets reveals notable gaps. First, current data collection and synthesis focus on data volume and lack consideration of fine-grained visual encodings and QA tasks, resulting in unbalanced data distribution divergent from practical CQA scenarios. Second, existing work follows the training recipe of the base MLLMs initially designed for natural images, under-exploring the adaptation to unique chart characteristics, such as rich text elements. To fill the gap, we propose a visualization-referenced instruction tuning approach to guide the training dataset enhancement and model development. Specifically, we propose a novel data engine to effectively filter diverse and high-quality data from existing datasets and subsequently refine and augment the data using LLM-based generation techniques to better align with practical QA tasks and visual encodings. Then, to facilitate the adaptation to chart characteristics, we utilize the enriched data to train an MLLM by unfreezing the vision encoder and incorporating a mixture-of-resolution adaptation strategy for enhanced fine-grained recognition. Experimental results validate the effectiveness of our approach. Even with fewer training examples, our model consistently outperforms state-of-the-art CQA models on established benchmarks. We also contribute a dataset split as a benchmark for future research. Source codes and datasets of this paper are available at https://github.com/zengxingchen/ChartQA-MLLM.
Xingchen Zeng, Haichuan Lin, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.4
2025 Latent Space Map for Visual Utilization of Generated Data
abstract
Samples produced by generative models, called Generated Samples (GSs), have become a critical supplement to those collected from the real world in data-centric applications. Domain experts typically randomly collect many GSs and manually select a few of interest for applications. However, the methodology lacks guidance to locate desirable ones that exhibit specific features or adhere to application-oriented metrics among infinite generable candidates. These samples are generally concentrated in a few small regions of the generative model's latent space, called Generative Latent Space (GLS). This paper presents Latent Space Map that projects a GLS onto a plane to help users locate regions rich in desirable GSs. Our research revolves around two challenges in constructing the map. First, many GSs in a GLS are low-quality and useless for applications. Excluding them from the projection is challenging for their irregular distribution. We employ a Monte Carlo-based method to capture a manifold for projection, where high-quality GSs are mainly distributed. Second, the GLS is high-dimensional and unbounded, complicating the projection. We design a manifold projection method that endows the map with desirable characteristics to achieve high display accuracy and effective pattern perception for users freely observing the manifold. We further develop a system integrating Latent Space Map to aid in GS selection and refinement. Real-world cases, quantitative experiments, and feedback from domain experts confirm the usability and effectiveness of our approach.
Jie Li 0006, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.3
2025 Antarctica storytelling: creating interactive story maps for polar regions with graphic-based approach
Liangwei Wang 0001, Zhan Wang 0001, Xi Zhao 0003, Fugee Tsung, Wei Zeng 0004
Vis. Comput.5
2025 A survey of visual insight mining: Connecting data and insights via visualization
abstract
Insight mining transforms complex data into actionable knowledge, enabling effective decision-making across diverse domains. Given the richness and interpretative power of visualizations, visual insight mining - the process of extracting meaningful insights from raw data through intuitive visual representations - has become increasingly vital. This survey systematically reviews the current landscape of visual insight mining, addressing the critical questions: “How can visualizations be generated from data?” and “How can insights be extracted from visualizations?” . Specifically, we delve into six distinct tasks (i.e., task decomposition, visualization generation, visualization recommendation, chart parsing, chart question answering, and insight generation) in the process of visual insight mining, and provide a comprehensive analysis of rule-based, learning-based, and large-model-based methods for each task. Based on the survey, we discuss current research challenges and outline future opportunities. By viewing visualization as a bridge in the data-to-insight path, this survey offers a structured foundation for further exploration in visual insight mining.
Yijie Lian, Jianing Hao, Wei Zeng 0004, Qiong Luo 0007
Vis. Informatics3
2025 EmotionLens: Interactive visual exploration of the circumplex emotion space in literary works via affective word clouds
abstract
Emotion (e.g., valence and arousal) is an important factor in literature (e.g., poetry and prose), and has rich values for plotting the life and knowledge of historical figures and appreciating the aesthetics of literary works. Currently, digital humanities and computational literature apply data statistics extensively in emotion analysis but lack visual analytics for efficient exploration. To fill the gap, we propose a user-centric approach that integrates advanced machine learning models and intuitive visualization for emotion analysis in literature. We make three main contributions. First, we consolidate a new emotion dataset of literary works in different periods, literary genres, and language contexts, augmented with fine-grained valence and arousal labels. Next, we design an interactive visual analytic system named EmotionLens , which allows users to perform multi-granularity (e.g., individual, group, society) and multi-faceted (e.g., distribution, chronology, correlation) analyses of literary emotions, supporting both exploratory and confirmatory approaches in digital humanities. Specifically, we introduce a novel affective word cloud with augmented word weight, position, and color, to facilitate literary text analysis from an emotional perspective. To validate the usability and effectiveness of EmotionLens , we provide two consecutive case studies, two user studies, and interviews with experts from different domains. Our results show that EmotionLens bridges literary text, emotion, and various other attributes, enables efficient knowledge discovery in massive data, and facilitates raising and validating domain-specific hypotheses in literature.
Bingyuan Wang, Wei Zeng 0004, Zeyu Wang 0003
Vis. Informatics5
2025 CineFolio: Cinematography-guided camera planning for immersive narrative visualization
abstract
Narrative visualization facilitates data presentation and communicates insights, while virtual reality can further enhance immersive and engaging experiences. The combination of these two research interests shows the potential to revolutionize the way data is presented and understood. Within the realm of narrative visualization, empirical evidence has particularly highlighted the importance of camera planning. However, existing works primarily rely on user-intensive manipulation of the camera, with little effort put into automating the process. To fill the gap, this paper proposes CineFolio , a semi-automated camera planning method to reduce manual effort and enhance user experience in immersive narrative visualization. CineFolio combines cinematic theories with graphics criteria, considering both information delivery and aesthetic enjoyment to ensure a comfortable and engaging experience. Specifically, we parametrize the considerations into optimizable camera properties and solve it as a constraint satisfaction problem (CSP) to realize common camera types for narrative visualization, namely overview camera for absorbing the scale, focus camera for detailed views, moving camera for animated transitions, and user-controlled camera allowing users to provide inputs to camera planning. We demonstrate the feasibility of our approach with cases of various data and chart types. To further evaluate our approach, we conducted a within-subject user study, comparing our automated method with manual camera control, and the results confirm both effectiveness of the guided navigation and expressiveness of the cinematic design for narrative visualization.
Zhan Wang 0001, Qian Zhu 0010, David Kei-Man Yip, Fugee Tsung, Wei Zeng 0004
Vis. Informatics5
2024 Make Interaction Situated: Designing User Acceptable Interaction for Situated Visualization in Public Environments
abstract
Situated visualization blends data into the real world to fulfill individuals’ contextual information needs. However, interacting with situated visualization in public environments faces challenges posed by users’ acceptance and contextual constraints. To explore appropriate interaction design, we first conduct a formative study to identify users’ needs for data and interaction. Informed by the findings, we summarize appropriate interaction modalities with eye-based, hand-based and spatially-aware object interaction for situated visualization in public environments. Then, through an iterative design process with six users, we explore and implement interactive techniques for activating and analyzing with situated visualization. To assess the effectiveness and acceptance of these interactions, we integrate them into an AR prototype and conduct a within-subjects study in public scenarios using conventional hand-only interactions as the baseline. The results show that participants preferred our prototype over the baseline, attributing their preference to the interactions being more acceptable, flexible, and practical in public.
Qian Zhu 0010, Wei Zeng 0004, Wai Tong, Weiyue Lin, Xiaojuan Ma
CHI3
2024 C2Ideas: Supporting Creative Interior Color Design Ideation with a Large Language Model
abstract
Interior color design is a creative process that endeavors to allocate colors to furniture and other elements within an interior space. While much research focuses on generating realistic interior designs, these automated approaches often misalign with user intention and disregard design rationales. Informed by a need-finding preliminary study, we develop C2Ideas, an innovative system for designers to creatively ideate color schemes enabled by an intent-aligned and domain-oriented large language model. C2Ideas integrates a three-stage process: Idea Prompting stage distills user intentions into color linguistic prompts; Word-Color Association stage transforms the prompts into semantically and stylistically coherent color schemes; and Interior Coloring stage assigns colors to interior elements complying with design principles. We also develop an interactive interface that enables flexible user refinement and interpretable reasoning. C2Ideas has undergone a series of indoor cases and user studies, demonstrating its effectiveness and high recognition of interactive functionality by designers.
Yihan Hou, Manling Yang, Jie Xu 0046, Wei Zeng 0004
CHI6
2024 PlantoGraphy: Incorporating Iterative Design Process into Generative Artificial Intelligence for Landscape Rendering
abstract
Landscape renderings are realistic images of landscape sites, allowing stakeholders to perceive better and evaluate design ideas. While recent advances in Generative Artificial Intelligence (GAI (generative artificial intelligence)) enable automated generation of landscape renderings, the End to End (endtoend) methods are not compatible with common design processes, leading to insufficient alignment with design idealizations and limited cohesion of iterative landscape design. Informed by a formative study for comprehending design requirements, we present PlantoGraphy, an iterative design system that allows for interactive configuration of generative artificial intelligence models to accommodate human-centered design practice. A two-stage pipeline is incorporated: first, the concretization module transforms conceptual ideas into concrete scene layouts with a domain-oriented large language model; and second, the illustration module converts scene layouts into realistic landscape renderings with a layout-guided diffusion model Fine-tune (finetune)ed through Low-Rank Adaptation (LoRA) (lora). PlantoGraphy has undergone a series of performance evaluations and user studies, demonstrating its effectiveness in landscape rendering generation and the high recognition of its interactive functionality.
Rong Huang 0007, Haichuan Lin, Chuanzhang Chen, Kang Zhang 0001, Wei Zeng 0004
CHI5
2024 VirtuWander: Enhancing Multi-modal Interaction for Virtual Tour Guidance through Large Language Models
abstract
Tour guidance in virtual museums encourages multi-modal interactions to boost user experiences, concerning engagement, immersion, and spatial awareness. Nevertheless, achieving the goal is challenging due to the complexity of comprehending diverse user needs and accommodating personalized user preferences. Informed by a formative study that characterizes guidance-seeking contexts, we establish a multi-modal interaction design framework for virtual tour guidance. We then design VirtuWander, a two-stage innovative system using domain-oriented large language models to transform user inquiries into diverse guidance-seeking contexts and facilitate multi-modal interactions. The feasibility and versatility of VirtuWander are demonstrated with virtual guiding examples that encompass various touring scenarios and cater to personalized preferences. We further evaluate VirtuWander through a user study within an immersive simulated museum. The results suggest that our system enhances engaging virtual tour experiences through personalized communication and knowledgeable assistance, indicating its potential for expanding into real-world scenarios.
Zhan Wang 0001, Linping Yuan, Liangwei Wang 0001, Bingchuan Jiang, Wei Zeng 0004
CHI5
2024 TypeDance: Creating Semantic Typographic Logos from Image through Personalized Generation
abstract
Semantic typographic logos harmoniously blend typeface and imagery to represent semantic concepts while maintaining legibility. Conventional methods using spatial composition and shape substitution are hindered by the conflicting requirement for achieving seamless spatial fusion between geometrically dissimilar typefaces and semantics. While recent advances made AI generation of semantic typography possible, the end-to-end approaches exclude designer involvement and disregard personalized design. This paper presents TypeDance, an AI-assisted tool incorporating design rationales with the generative model for personalized semantic typographic logo design. It leverages combinable design priors extracted from uploaded image exemplars and supports type-imagery mapping at various structural granularity, achieving diverse aesthetic designs with flexible control. Additionally, we instantiate a comprehensive design workflow in TypeDance, including ideation, selection, generation, evaluation, and iteration. A two-task user evaluation, including imitation and creation, confirmed the usability of TypeDance in design across different usage scenarios.
Shishi Xiao, Liangwei Wang 0001, Xiaojuan Ma, Wei Zeng 0004
CHI4
2024 IntentTuner: An Interactive Framework for Integrating Human Intentions in Fine-tuning Text-to-Image Generative Models
abstract
Fine-tuning facilitates the adaptation of text-to-image generative models to novel concepts (e.g., styles and portraits), empowering users to forge creatively customized content. Recent efforts on fine-tuning focus on reducing training data and lightening computation overload but neglect alignment with user intentions, particularly in manual curation of multi-modal training data and intent-oriented evaluation. Informed by a formative study with fine-tuning practitioners for comprehending user intentions, we propose IntentTuner, an interactive framework that intelligently incorporates human intentions throughout each phase of the fine-tuning workflow. IntentTuner enables users to articulate training intentions with imagery exemplars and textual descriptions, automatically converting them into effective data augmentation strategies. Furthermore, IntentTuner introduces novel metrics to measure user intent alignment, allowing intent-aware monitoring and evaluation of model training. Application exemplars and user studies demonstrate that IntentTuner streamlines fine-tuning, reducing cognitive effort and yielding superior models compared to the common baseline tool.
Xingchen Zeng, Ziyao Gao, Wei Zeng 0004
CHI4
2024 GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
abstract
This work tackles the problem of geo-localization with a new paradigm using a large vision-language model (LVLM) augmented with human inference knowledge. A primary challenge here is the scarcity of data for training the LVLM - existing street-view datasets often contain numerous low-quality images lacking visual clues, and lack any reasoning inference. To address the data-quality issue, we devise a CLIP-based network to quantify the degree of street-view images being locatable, leading to the creation of a new dataset comprising highly locatable street views. To enhance reasoning inference, we integrate external knowledge obtained from real geo-localization games, tapping into valuable human inference capabilities. The data are utilized to train GeoReasoner, which undergoes fine-tuning through dedicated reasoning and location-tuning stages. Qualitative and quantitative evaluations illustrate that GeoReasoner outperforms counterpart LVLMs by more than 25% at country-level and 38% at city-level geo-localization tasks, and surpasses StreetCLIP performance while requiring fewer training resources. The data and code are available at https://github.com/lingli1996/GeoReasoner.
Yu Ye 0002, Bingchuan Jiang, Wei Zeng 0004
ICML4
2024 A-Map: Interactive Visual Exploration of Intercity Accessibility Dynamics Based on Railway Network Data
abstract
Railway transportation is closely linked to everyday lives while also aiding domain experts in analyzing national or regional development. However, discrepancies between travel time reduced by railways and actual distance pose challenges in visualizing the accessibility. Existing methods either struggle to simultaneously depict the accessibility relationships among all cities or disregard genuine geographical positions, leading to spatial cognitive confusion. In this work, we propose a novel approach to balance between the geographical positions and travel time between cities. We construct linear cartograms to represent accessibility while preserving better local railway network structures compared with existing methods. We further propose the A-Map system, enabling users to interactively explore the extensive development of China’s railway system over decades. We validate the effectiveness of our proposed layout algorithm from qualitative metric evaluation and illustrate the applicability of our system with results.
Kaichen Nie, Hanning Shao, Yuchu Luo, Min Tian 0008, Wei Zeng 0004, Xiaoru Yuan
PacificVis6
2024 HarmonyWave: Immersive Space Sound Therapy with Audio-Visual Synesthesia
abstract
The paper presents a study on sound therapy installation HarmonyWave utilizing audio-visual synesthesia to alleviate anxiety. At the core of HarmonyWave lies the mapping of sounds to visuals with three distinct methods: laser vibration for line-based effects, water vibration for color-based effects, and water vibration for texture-based effects. On the basis, we design an interactive installation for an immersive space sound therapy experience within 3D space. The setup enables the audience to simultaneously hear the sound and observe superimposed images while actively engaging with the interactive elements. HarmonyWave was implemented with 20 natural sounds, specifically focusing on white noise known for inducing a calm and soothing emotional state in urban environments. We envision HarmonyWave as a tool to alleviate tension and anxiety in urban settings, serving a therapeutic function. We also delve into the possibilities and diversity offered by the three sensory mediums for mapping audio to visuals. © 2024 Copyright held by the owner/author(s).
Wei Zeng 0004
VINCI3
2024 Generating Virtual Reality Stroke Gesture Data from Out-of-Distribution Desktop Stroke Gesture Data
abstract
This paper exploits ubiquitous desktop interaction data as an input source for generating virtual reality (VR) interaction data, which can benefit tasks like user behavior analysis and experience enhancement. Time-varying stroke gestures are selected as the primary focus because of their prevalence across various applications and their diverse patterns. The commonalities (e.g., features like velocity and curvature) between desktop and VR strokes allow the generation of additional dimensions (e.g., z vectors) in VR strokes. However, distribution shifts exist between different interaction environments (i.e., desktop vs. VR), and within the same interaction environment for different strokes by various users, making it challenging to build models capable of generalizing to unseen distributions. To address the challenges, we formulate the problem of generating VR strokes from desktop strokes as a conditional time series generation problem, aiming to learn representations that are capable of handling out-of-distribution data. We propose a novel architecture based on conditional generative adversarial networks, with the generator encompassing three steps: discretizing the output space, characterizing latent distributions, and learning conditional domain-invariant representations. We evaluate the effectiveness of our methods by comparing them with state-of-the-art time series generation models and conducting ablation studies. We further illustrate the applicability of the enriched VR datasets through two applications: VR stroke classification and stroke prediction.
Linping Yuan, Boyu Li 0007, Jindong Wang 0001, Huamin Qu, Wei Zeng 0004
VR5
2024 HoLens: A visual analytics design for higher-order movement modeling and visualization
abstract
Higher-order patterns reveal sequential multistep state transitions, which are usually superior to origin-destination analyses that depict only first-order geospatial movement patterns. Conventional methods for higher-order movement modeling first construct a directed acyclic graph (DAG) of movements and then extract higher-order patterns from the DAG. However, DAG-based methods rely heavily on identifying movement keypoints, which are challenging for sparse movements and fail to consider the temporal variants critical for movements in urban environments. To overcome these limitations, we propose HoLens, a novel approach for modeling and visualizing higher-order movement patterns in the context of an urban environment. HoLens mainly makes twofold contributions: First, we designed an auto-adaptive movement aggregation algorithm that self-organizes movements hierarchically by considering spatial proximity, contextual information, and temporal variability. Second, we developed an interactive visual analytics interface comprising well-established visualization techniques, including the H-Flow for visualizing the higher-order patterns on the map and the higher-order state sequence chart for representing the higher-order state transitions. Two real-world case studies demonstrate that the method can adaptively aggregate data and exhibit the process of exploring higher-order patterns using HoLens. We also demonstrate the feasibility, usability, and effectiveness of our approach through expert interviews with three domain experts.
Zezheng Feng, Hongjun Wang 0007, Jianing Hao, Shuang-Hua Yang, Wei Zeng 0004, Huamin Qu
Comput. Vis. Media6
2024 The Contemporary Art of Image Search: Iterative User Intent Expansion via Vision-Language Model
abstract
Image search is an essential and user-friendly method to explore vast galleries of digital images. However, existing image search methods heavily rely on proximity measurements like tag matching or image similarity, requiring precise user inputs for satisfactory results. To meet the growing demand for a contemporary image search engine that enables accurate comprehension of users' search intentions, we introduce an innovative user intent expansion framework. Our framework leverages visual-language models to parse and compose multi-modal user inputs to provide more accurate and satisfying results. It comprises two-stage processes: 1) a parsing stage that incorporates a language parsing module with large language models to enhance the comprehension of textual inputs, along with a visual parsing module that integrates an interactive segmentation module to swiftly identify detailed visual elements within images; and 2) a logic composition stage that combines multiple user search intents into a unified logic expression for more sophisticated operations in complex searching scenarios. Moreover, the intent expansion framework enables users to perform flexible contextualized interactions with the search results to further specify or adjust their detailed search intents iteratively. We implemented the framework into an image search system for NFT (non-fungible token) search and conducted a user study to evaluate its usability and novel properties. The results indicate that the proposed framework significantly improves users' image search experience. Particularly the parsing and contextualized interactions prove useful in allowing users to express their search intents more accurately and engage in a more enjoyable iterative search experience.
Qian Zhu 0010, Shishi Xiao, Kang Zhang 0001, Wei Zeng 0004
Proc. ACM Hum. Comput. Interact.5
2024 MetroBUX: A Topology-Based Visual Analytics for Bus Operational Uncertainty EXploration
abstract
In the public transportation system, punctuality benefits both bus operation and passengers’ travel experience. However, uncertainty exists due to complex traffic conditions and heterogeneous driving behaviors. To analyze bus operational uncertainty, transport planners and bus operators need a tool that supports multi-granular modeling, spatio-temporal representation, and interactive exploration. To meet the requirement, we present MetroBUX, a visual analytics system for$B$us operational$U$ncertainty e$X$ploration. MetroBUX aligns daily bus trips and models stop-level uncertainty of bus arrival time. It has a consolidated interface with three main views: Map View for presenting the spatial distribution of uncertainty, Temporal View for tracking the evolution of uncertainty, and Trip View for inspecting uncertainty propagation. Specifically, MetroBUX enables integrated spatio-temporal analysis by connecting topological uncertainty distribution at different periods in a nested tracking graph. Furthermore, it supports interactive and hierarchical exploration, including region-, route-, trip-, and stop-level analysis. Case studies on real-world bus operational data and domain experts’ feedback demonstrate the efficiency of MetroBUX.
Shishi Xiao, Lingdan Shao, Bo Du 0004, Yang Wang 0006, Qiaomu Shen, Wei Zeng 0004
IEEE Trans. Intell. Transp. Syst.7
2024 $\mathsf {CheetahTraj}$CheetahTraj: Efficient Visualization for Large Trajectory Dataset With Quality Guarantee
abstract
Visualizing large-scale trajectory dataset is a core subroutine for many applications. However, rendering all trajectories could result in severe visual clutter and incur long visualization delays due to large data volume. Naively sampling the trajectories reduces visualization time but usually harms visual quality, i.e., the generated visualizations may look substantially different from the exact ones without sampling. In this paper, we propose$\mathsf {CheetahTraj}$, a principled sampling framework that achieves both high visualization quality and low visualization latency. We first define thevisual quality functionmeasuring the similarity between two visualizations, based on which we formulate the quality optimal sampling problem (${\sf QOSP}$). To solve${\sf QOSP}$, we design theVisualQualityGuaranteedSampling algorithms, which reduce visual clutter while guaranteeing visual quality by considering both trajectory data distribution and human perception properties. We also develop a quad-tree-based index ($\mathsf {InvQuad}$) that allows using trajectory samples computed offline for interactive online visualization. Extensive experiments including case-, user-, and quantitative-studies are conducted on three real-world trajectory datasets, and the results show that$\mathsf {CheetahTraj}$consistently provides higher visual quality and better efficiency than baseline methods. Compared with visualizing all trajectories,$\mathsf {CheetahTraj}$reduces the visualization latency by up to 3 orders of magnitude while avoiding visual clutter.
Qiaomu Shen, Chaozu Zhang, Xiao Yan 0002, Dan Zeng 0002, Wei Zeng 0004, Bo Tang 0016
IEEE Trans. Knowl. Data Eng.6
2024 : Diagnosing Time Representations for Time-Series Forecasting with Counterfactual Explanations
abstract
Deep learning (DL) approaches are being increasingly used for time-series forecasting, with many efforts devoted to designing complex DL models. Recent studies have shown that the DL success is often attributed to effective data representations, fostering the fields of feature engineering and representation learning. However, automated approaches for feature learning are typically limited with respect to incorporating prior knowledge, identifying interactions among variables, and choosing evaluation metrics to ensure that the models are reliable. To improve on these limitations, this paper contributes a novel visual analytics framework, namely TimeTuner, designed to help analysts understand how model behaviors are associated with localized correlations, stationarity, and granularity of time-series representations. The system mainly consists of the following two-stage technique: We first leverage counterfactual explanations to connect the relationships among time-series representations, multivariate features and model predictions. Next, we design multiple coordinated views including a partition-based correlation matrix and juxtaposed bivariate stripes, and provide a set of interactions that allow users to step into the transformation selection process, navigate through the feature space, and reason the model performance. We instantiate TimeTuner with two transformation methods of smoothing and sampling, and demonstrate its applicability on real-world time-series forecasting of univariate sunspots and multivariate air pollutants. Feedback from domain experts indicates that our system can help characterize time-series representations and guide the feature engineering processes.
Jianing Hao, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.4
2024 Let the Chart Spark: Embedding Semantic Context into Chart with Text-to-Image Generative Model
abstract
Pictorial visualization seamlessly integrates data and semantic context into visual representation, conveying complex information in an engaging and informative manner. Extensive studies have been devoted to developing authoring tools to simplify the creation of pictorial visualizations. However, mainstream works follow a retrieving-and-editing pipeline that heavily relies on retrieved visual elements from a dedicated corpus, which often compromise data integrity. Text-guided generation methods are emerging, but may have limited applicability due to their predefined entities. In this work, we propose ChartSpark, a novel system that embeds semantic context into chart based on text-to-image generative models. ChartSpark generates pictorial visualizations conditioned on both semantic context conveyed in textual inputs and data information embedded in plain charts. The method is generic for both foreground and background pictorial generation, satisfying the design practices identified from empirical research into existing pictorial visualizations. We further develop an interactive visual interface that integrates a text analyzer, editing module, and evaluation module to enable users to generate, modify, and assess pictorial visualizations. We experimentally demonstrate the usability of our tool, and conclude with a discussion of the potential of using text-to-image generative models combined with an interactive interface for visualization design.
Shishi Xiao, Suizi Huang, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.5
2024 VISAtlas: An Image-Based Exploration and Query System for Large Visualization Collections via Neural Image Embedding
abstract
High-quality visualization collections are beneficial for a variety of applications including visualization reference and data-driven visualization design. The visualization community has created many visualization collections, and developed interactive exploration systems for the collections. However, the systems are mainly based on extrinsic attributes like authors and publication years, whilst neglect intrinsic property (i.e., visual appearance) of visualizations, hindering visual comparison and query of visualization designs. This paper presents VISAtlas, an image-based approach empowered by neural image embedding, to facilitate exploration and query for visualization collections. To improve embedding accuracy, we create a comprehensive collection of synthetic and real-world visualizations, and use it to train a convolutional neural network (CNN) model with a triplet loss for taxonomical classification of visualizations. Next, we design a coordinated multiple view (CMV) system that enables multi-perspective exploration and design retrieval based on visualization embeddings. Specifically, we design a novel embedding overview that leverages contextual layout framework to preserve the context of the embedding vectors with the associated visualization taxonomies, and density plot and sampling techniques to address the overdrawing problem. We demonstrate in three case studies and one user study the effectiveness of VISAtlas in supporting comparative analysis of visualization collections, exploration of composite visualizations, and image-based retrieval of visualization designs. The studies reveal that real-world visualization collections (e.g., Beagle and VIS30K) better accord with the richness and diversity of visualization designs than synthetic collections (e.g., Data2Vis), inspiring composite visualizations are identified in real-world collections, and distinct design patterns exist in visualizations from different sources.
Rong Huang 0007, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.3
2024 Semi-Automatic Layout Adaptation for Responsive Multiple-View Visualization Design
abstract
Multiple-view (MV) visualizations have become ubiquitous for visual communication and exploratory data visualization. However, most existing MV visualizations are designed for the desktop, which can be unsuitable for the continuously evolving displays of varying screen sizes. In this article, we present a two-stage adaptation framework that supports the automated retargeting and semi-automated tailoring of a desktop MV visualization for rendering on devices with displays of varying sizes. First, we cast layout retargeting as an optimization problem and propose a simulated annealing technique that can automatically preserve the layout of multiple views. Second, we enable fine-tuning for the visual appearance of each view, using a rule-based auto configuration method complemented with an interactive interface for chart-oriented encoding adjustment. To demonstrate the feasibility and expressivity of our proposed approach, we present a gallery of MV visualizations that have been adapted from the desktop to small displays. We also report the result of a user study comparing visualizations generated using our approach with those by existing methods. The outcome indicates that the participants generally prefer visualizations generated using our approach and find them to be easier to use.
Wei Zeng 0004, Xi Chen 0072, Yihan Hou, Lingdan Shao, Zhe Chu, Remco Chang
IEEE Trans. Vis. Comput. Graph.1
2024 Generative AI for visualization: State of the art and future directions
abstract
Generative AI (GenAI) has witnessed remarkable progress in recent years and demonstrated impressive performance in various generation tasks in different domains such as computer vision and computational design. Many researchers have attempted to integrate GenAI into visualization framework, leveraging the superior generative capacity for different operations. Concurrently, recent major breakthroughs in GenAI like diffusion model and large language model have also drastically increase the potential of GenAI4VIS. From a technical perspective, this paper looks back on previous visualization studies leveraging GenAI and discusses the challenges and opportunities for future research. Specifically, we cover the applications of different types of GenAI methods including sequence, tabular, spatial and graph generation techniques for different tasks of visualization which we summarize into four major stages: data enhancement, visual mapping generation, stylization and interaction. For each specific visualization sub-task, we illustrate the typical data and concrete GenAI algorithms, aiming to provide in-depth understanding of the state-of-the-art GenAI4VIS techniques and their limitations. Furthermore, based on the survey, we discuss three major aspects of challenges and research opportunities including evaluation, dataset, and the gap between end-to-end GenAI methods and visualizations. By summarizing different generation algorithms, their current applications and limitations, this paper endeavors to provide useful insights for future GenAI4VIS research.
Jianing Hao, Yihan Hou, Zhan Wang 0001, Shishi Xiao, Yuyu Luo, Wei Zeng 0004
Vis. Informatics7
2023 NFTeller: Dual-centric Visual Analytics for Assessing Market Performance of NFT Collectibles
abstract
Non-fungible tokens (NFTs) have recently gained widespread popularity as an alternative investment. However, the lack of assessment criteria has caused intense volatility in NFT marketplaces. Identifying attributes impacting the market performance of NFT collectibles is crucial but challenging due to the massive amount of heterogeneous and multi-modal data in NFT transactions, e.g., social media texts, numerical trading data, and images. To address this challenge, we introduce an interactive dual-centric visual analytics system, NFTeller, to facilitate users’ analysis. First, we collaborate with five domain experts to distill static and dynamic impact attributes and collect relevant data. Next, we derive six analysis tasks and develop NFTeller to present the evolution of NFT transactions and correlate NFTs’ market performance with impact attributes. Notably, we create an augmented chord diagram with a radial stacked bar chart to explore intersections between NFT collection projects and whale accounts. Finally, we conduct three case studies and interview domain experts to evaluate the effectiveness and usability of this system. As such, we gain in-depth insights into assessing NFT collectibles and detecting opportune moments for investment.
Yifan Cao 0001, Meng Xia 0002, Kento Shigyo, Furui Cheng, Qianhang Yu, Xingxing Yang 0005, Yang Wang 0020, Wei Zeng 0004, Huamin Qu
VINCI8
2023 The Rich, the Poor, and the Ugly: An Aesthetic-Perspective Assessment of NFT Values
abstract
The adoption of non-fungible tokens (NFTs) has revolutionized digital art transactions, providing artists with unprecedented opportunities to tokenize and monetize their generative creations, leading to increased scrutiny and demand within blockchain-oriented marketplaces. The pricing of NFT artworks, however, exhibits substantial variations within and across collections, influenced by various factors. This study aims to investigate the relationship between visual features and pricing, shedding light on the variations underlying the pricing of NFTs. First, measures of both computational aesthetics and visual complexity were applied to extract multi-faceted visual aesthetic features, encompassing aesthetic factors such as color and composition as well as complexity factors like entropy. Second, with extracted visual aesthetic features and preprocessed price data, the study proceeds to conduct correlation analysis within collections and statistical modeling across collections. Through these approaches, we reveal a moderate correlation between visual features and prices within collections, while also identifying different influential visual features across collections. The differential performance of price models highlights the distinctiveness and unique pricing characteristics of NFT collections.
Yihan Chen 0006, Wei Zeng 0004
VINCI3
2023 LOOP Meditation: Enhancing Novice's VR Meditation Experience with Physical Movement
abstract
Virtual reality (VR) and associated technologies have rapidly grew, creating new opportunities for improving mental health. In order to provide an immersive and concentrated meditation experience, this paper offers the idea of VR-assisted meditation, which integrates the advantages of VR technology with mindfulness techniques. The suggested technique, which is known as LOOP Meditation, is mainly aimed toward novice meditators and places a strong emphasis on the value of movement and breath awareness when meditating. Existing VR experiences and meditation applications sometimes ignore the value of including physical movement, which can improve mindful body sensations and maintain interest. By creating a software that incorporates body movement and breath sensing into virtual reality surroundings, LOOP Meditation addresses this gap. The LOOP Meditation design and implementation are examined in this paper along with its possible advantages and consequences for people looking to develop their meditation practice. A pilot research comparing the efficacy of this strategy to conventional meditation techniques is offered. The study’s findings add to the expanding body of knowledge on VR-assisted meditation and demonstrate how it may have a positive effect on mental health.
Shihan Fu, Liangliang Qiang, Wei Zeng 0004
VINCI3
2023 Does Where You are Matter? A Visual Analytics System for COVID-19 Transmission Based on Social Hierarchical Perspective
abstract
The COVID-19 pandemic requires multidisciplinary efforts to address its profound social and economic repercussions. Combining social hierarchical perspectives and a Multiple Coordinated View (MCV) visualization system, this paper depicts how social and physical residential environment shapes individuals’ infection risk during such pandemic. Through analyzing the travel records of 8000+ confirmed cases in spatial and temporal channels, we identify that there exists segregation of virus transmission among different social classes and individuals from deprived neighborhoods exhibit a higher risk of the virus infection. Leveraging our proposed interactive visualization system, policymakers and stakeholders can make more informed decisions to effectively manage and contain the spread of infectious pandemics like COVID-19.
Jianing Hao, Xibin Jiang, Wei Zeng 0004
VINCI4
2023 Storytelling in Frozen Frontier: Exploring Graphic-Based Approach for Creating Interactive Story Maps in Antarctica
abstract
Story maps have been widely utilized to provide a visual and spatial framework for storytelling. However, existing story map tools have limitations in creating diverse narrative structures and providing interactive options, and cannot effectively render maps for polar regions due to tile-based mapping constraints. In this paper, we propose a graphic-based approach to overcome these challenges and develop a workflow for creating story maps specifically designed for polar regions. A primary contribution is to provide heuristic strategies for story map design and explore the potential of story maps in visualizing and disseminating polar culture. We summarize the main design tasks involved in story map creation and introduce three map-based visual narrative strategies, i.e., attention cue, linkage of map and other visual elements, and cartographic interaction. Additionally, we delve into the importance of storyboard design, taking into account logic, time order, and map granularity. To demonstrate the effectiveness of our proposed story map design method, we create story map cases focusing on the exploration history of Antarctica. These cases showcase the diverse and interactive nature of the story maps created using our approach. We discuss the limitations and challenges in creating story maps and identify potential research opportunities from our study.
Liangwei Wang 0001, Zhan Wang 0001, Xi Zhao 0003, Wei Zeng 0004
VINCI4
2023 WYTIWYR: A User Intent-Aware Framework with Multi-modal Inputs for Visualization Retrieval
abstract
Abstract Retrieving charts from a large corpus is a fundamental task that can benefit numerous applications such as visualization recommendations. The retrieved results are expected to conform to both explicit visual attributes (e.g., chart type, colormap) and implicit user intents (e.g., design style, context information) that vary upon application scenarios. However, existing example‐based chart retrieval methods are built upon non‐decoupled and low‐level visual features that are hard to interpret, while definition‐based ones are constrained to pre‐defined attributes that are hard to extend. In this work, we propose a new framework, namelyWYTIWYR (What‐You‐Think‐Is‐What‐You‐Retrieve), that integrates user intents into the chart retrieval process. The framework consists of two stages: first, theAnnotationstage disentangles the visual attributes within the query chart; and second, theRetrievalstage embeds the user's intent with customized text prompt as well as bitmap query chart, to recall targeted retrieval result. We develop aprototypeWYTIWYRsystem leveraging a contrastive language‐image pre‐training (CLIP) model to achieve zero‐shot classification as well as multi‐modal input encoding, and test the prototype on a large corpus with charts crawled from the Internet. Quantitative experiments, case studies, and qualitative interviews are conducted. The results demonstrate the usability and effectiveness of our proposed framework.
Shishi Xiao, Yihan Hou, Cheng Jin 0003, Wei Zeng 0004
Comput. Graph. Forum4
2023 Modeling Spatial Nonstationarity via Deformable Convolutions for Deep Traffic Flow Prediction
abstract
Deep neural networks are being increasingly used for short-term traffic flow prediction, which can be generally categorized as CNNs or GNNs. CNNs typically partition an underlying territory into grid-like spatial units, and employ standard convolutions to learn spatial dependence among the units. However, standard convolutions with fixed geometric structures cannot fully model the nonstationary characteristics of local traffic flows. To overcome the deficiency, we introduce deformable convolution that augments the spatial sampling locations with additional offsets, to enhance the modeling capability of spatial nonstationarity. We design a deep deformable convolutional residual network, namely DeFlow-Net, that can effectively model global spatial dependence, local spatial nonstationarity, and temporal periodicity of traffic flows. Furthermore, to better fit with convolutions, we suggest to first aggregate traffic flows according to pre-conceived regions or self-organized regions based on traffic flows, then dispose to sequentially organized raster images for network input. Extensive experiments on real-world traffic flows demonstrate that DeFlow-Net outperforms GNNs and existing CNNs using standard convolutions, and spatial partition by pre-conceived regions or self-organized regions further enhances the performance. We also demonstrate the advantage of DeFlow-Net in maintaining spatial autocorrelation, and reveal the impacts of partition shapes and scales on deep traffic flow prediction.
Wei Zeng 0004, Chengqiao Lin, Kang Liu 0010, Juncong Lin, Anthony K. H. Tung
IEEE Trans. Knowl. Data Eng.1
2023 Creative and Progressive Interior Color Design with Eye-tracked User Preference
abstract
Interior scene colorization is vastly demanded in areas such as personalized architecture design. Existing works either require manual efforts to colorize individual objects or conform to fixed color patterns automatically learned from prior knowledge, whilst neglecting user preference. Quantitatively identifying user preferences is challenging, particularly at the early stage of the design process. The 3D setup also presents new challenges as the inhabitant can observe from any possible viewpoint. We propose a representative view selection method based on visual attention and a progressive preference inference model. We particularly focus on the progressive integration of eye-tracked user preference, which enables the assistance in creativity support and allows the possibility of convergent thinking. A series of user studies have been conducted to validate the effectiveness of the proposed view selection method, preference inference model and the creativity support mechanism.
Shihui Guo, Yubin Shi, Pintong Xiao, Yinan Fu, Juncong Lin, Wei Zeng 0004, Tong-Yee Lee
ACM Trans. Comput. Hum. Interact.6
2023 ActFloor-GAN: Activity-Guided Adversarial Networks for Human-Centric Floorplan Design
abstract
We present a novel two-stage approach for automated floorplan design in residential buildings with a given exterior wall boundary. Our approach has the unique advantage of being human-centric, that is, the generated floorplans can be geometrically plausible, as well as topologically reasonable to enhance resident interaction with the environment. From the input boundary, we first synthesize a human-activity map that reflects both the spatial configuration and human-environment interaction in an architectural space. We propose to produce the human-activity map either automatically by a pre-trained generative adversarial network (GAN) model, or semi-automatically by synthesizing it with user manipulation of the furniture. Second, we feed the human-activity map into our deep framework ActFloor-GAN to guide a pixel-wise prediction of room types. We adopt a re-formulated cycle-consistency constraint in ActFloor-GAN to maximize the overall prediction performance, so that we can produce high-quality room layouts that are readily convertible to vectorized floorplans. Experimental results show several benefits of our approach. First, a quantitative comparison with prior methods shows superior performance of leveraging the human-activity map in predicting piecewise room types. Second, a subjective evaluation by architects shows that our results have compelling quality as professionally-designed floorplans and much better than those generated by existing methods in terms of the room layout topology. Last, our approach allows manipulating the furniture placement, considers the human activities in the environment, and enables the incorporation of user-design preferences.
Wei Zeng 0004, Xi Chen 0072, Yu Ye 0002, Yu Qiao 0001, Chi-Wing Fu
IEEE Trans. Vis. Comput. Graph.2
2023 Effects of View Layout on Situated Analytics for Multiple-View Representations in Immersive Visualization
abstract
Multiple-view (MV) representations enabling multi-perspective exploration of large and complex data are often employed on 2D displays. The technique also shows great potential in addressing complex analytic tasks in immersive visualization. However, although useful, the design space of MV representations in immersive visualization lacks in deep exploration. In this paper, we propose a new perspective to this line of research, by examining the effects of view layout for MV representations on situated analytics. Specifically, we disentangle situated analytics in perspectives of situatedness regarding spatial relationship between visual representations and physical referents, and analytics regarding cross-view data analysis including filtering, refocusing, and connecting tasks. Through an in-depth analysis of existing layout paradigms, we summarize design trade-offs for achieving high situatedness and effective analytics simultaneously. We then distill a list of design requirements for a desired layout that balances situatedness and analytics, and develop a prototype system with an automatic layout adaptation method to fulfill the requirements. The method mainly includes a cylindrical paradigm for egocentric reference frame, and a force-directed method for proper view-view, view-user, and view-referent proximities and high view visibility. We conducted a formal user study that compares layouts by our method with linked and embedded layouts. Quantitative results show that participants finished filtering- and connecting-centered tasks significantly faster with our layouts, and user feedback confirms high usability of the prototype system.
Zhen Wen 0001, Wei Zeng 0004, Luoxuan Weng, Mingliang Xu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.2
2022 Saliency-aware color harmony models for outdoor signboard
Yanna Lin, Wei Zeng 0004, Yu Ye 0002, Huamin Qu
Comput. Graph.2
2022 Large-Scale Urban Multiple-Modal Transport Evacuation Model for Mass Gathering Events Considering Pedestrian and Public Transit System
abstract
Mass gathering events occur frequently in urban regions. Not only serious traffic jams, but also safety risks are consequently caused. Although many evacuation strategies have been proposed, the spatiotemporal coordination issue of multiple-modal transport tools is not solved well to deal with the efficiency and safety risk during the traditional evacuations. This study presented a large-scale multi-modal transport macro-optimization evacuation model for urban mass gathering events. By taking full advantages of each kind of public transit vehicles and optimizing their spatiotemporal cooperation with pedestrian evacuation, our models can greatly improve the global evacuation efficiency. Experiments on a realistic event was carried to validate the proposed model. The numerical results demonstrated that, under the premise of no increasing additional vehicle supply, the efficiency of proposed multi-modal transport evacuation can be improved by 49.7%, 46.5%, 118.7%, and 20.8%, compared with pure metro-, bus-, taxi- single-modal-transport evacuation, and simulated multi-modal transport evacuation without adequate optimization, respectively.
Jincheng Jiang, Wei Tu 0001, Hui Kong 0006, Wei Zeng 0004, Milan Konecný
IEEE Trans. Intell. Transp. Syst.4
2022 Deep Colormap Extraction From Visualizations
abstract
This article presents a new approach based on deep learning to automatically extract colormaps from visualizations. After summarizing colors in an input visualization image as a Lab color histogram, we pass the histogram to a pre-trained deep neural network, which learns to predict the colormap that produces the visualization. To train the network, we create a new dataset of ∼ 64K visualizations that cover a wide variety of data distributions, chart types, and colormaps. The network adopts an atrous spatial pyramid pooling module to capture color features at multiple scales in the input color histograms. We then classify the predicted colormap as discrete or continuous, and refine the predicted colormap based on its color histogram. Quantitative comparisons to existing methods show the superior performance of our approach on both synthetic and real-world visualizations. We further demonstrate the utility of our method with two use cases, i.e., color transfer and color remapping.
Linping Yuan, Wei Zeng 0004, Siwei Fu, Zhiliang Zeng, Haotian Li 0001, Chi-Wing Fu, Huamin Qu
IEEE Trans. Vis. Comput. Graph.2
2021 UrbanVR: An immersive analytics system for context-aware urban design
abstract
Urban design is a highly visual discipline that requires visualization for informed decision making. However, traditional urban design tools are mostly limited to representations on 2D displays that lack intuitive awareness. The popularity of head-mounted displays (HMDs) promotes a promising alternative with consumer-grade 3D displays. We introduce UrbanVR, an immersive analytics system with effective visualization and interaction techniques, to enable architects to assess designs in a virtual reality (VR) environment. Specifically, UrbanVR incorporates 1) a customized parallel coordinates plot (PCP) design to facilitate quantitative assessment of high-dimensional design metrics, 2) a series of egocentric interactions, including gesture interactions and handle-bar metaphors, to facilitate user interactions, and 3) a viewpoint optimization algorithm to help users explore both the PCP for quantitative analysis, and objects of interest for context awareness. Effectiveness and feasibility of the system are validated through quantitative user studies and qualitative expert feedbacks.
Wei Zeng 0004
Comput. Graph.2
2021 FloorLevel-Net: Recognizing Floor-Level Lines With Height-Attention-Guided Multi-Task Learning
abstract
The ability to recognize the position and order of the floor-level lines that divide adjacent building floors can benefit many applications, for example, urban augmented reality (AR). This work tackles the problem of locating floor-level lines in street-view images, using a supervised deep learning approach. Unfortunately, very little data is available for training such a network - current street-view datasets contain either semantic annotations that lack geometric attributes, or rectified facades without perspective priors. To address this issue, we first compile a new dataset and develop a new data augmentation scheme to synthesize training samples by harassing (i) the rich semantics of existing rectified facades and (ii) perspective priors of buildings in diverse street views. Next, we design FloorLevel-Net, a multi-task learning network that associates explicit features of building facades and implicit floor-level lines, along with a height-attention mechanism to help enforce a vertical ordering of floor-level lines. The generated segmentations are then passed to a second-stage geometry post-processing to exploit self-constrained geometric priors for plausible and consistent reconstruction of floor-level lines. Quantitative and qualitative evaluations conducted on assorted facades in existing datasets and street views from Google demonstrate the effectiveness of our approach. Also, we present context-aware image overlay results and show the potentials of our approach in enriching AR-related applications. Project website: https://wumengyangok.github.io/Project/FloorLevelNet.
Mengyang Wu, Wei Zeng 0004, Chi-Wing Fu
IEEE Trans. Image Process.2
2021 Composition and Configuration Patterns in Multiple-View Visualizations
abstract
Multiple-view visualization (MV) is a layout design technique often employed to help users see a large number of data attributes and values in a single cohesive representation. Because of its generalizability, the MV design has been widely adopted by the visualization community to help users examine and interact with large, complex, and high-dimensional data. However, although ubiquitous, there has been little work to categorize and analyze MVs in order to better understand its design space. As a result, there has been little to no guideline in how to use the MV design effectively. In this paper, we present an in-depth study of how MVs are designed in practice. We focus on two fundamental measures of multiple-view patterns: composition, which quantifies what view types and how many are there; and configuration, which characterizes spatial arrangement of view layouts in the display space. We build a new dataset containing 360 images of MVs collected from IEEE VIS, EuroVis, and PacificVis publications 2011 to 2019, and make fine-grained annotations of view types and layouts for these visualization images. From this data we conduct composition and configuration analyses using quantitative metrics of term frequency and layout topology. We identify common practices around MVs, including relationship of view types, popular view layouts, and correlation between view types and layouts. We combine the findings into a MV recommendation system, providing interactive tools to explore the design space, and support example-based design.
Xi Chen 0072, Wei Zeng 0004, Yanna Lin, Hayder Al-Maneea, Jonathan Roberts 0002, Remco Chang
IEEE Trans. Vis. Comput. Graph.2
2021 Topology Density Map for Urban Data Visualization and Analysis
abstract
Density map is an effective visualization technique for depicting the scalar field distribution in 2D space. Conventional methods for constructing density maps are mainly based on Euclidean distance, limiting their applicability in urban analysis that shall consider road network and urban traffic. In this work, we propose a new method named Topology Density Map, targeting for accurate and intuitive density maps in the context of urban environment. Based on the various constraints of road connections and traffic conditions, the method first constructs a directed acyclic graph (DAG) that propagates nonlinear scalar fields along 1D road networks. Next, the method extends the scalar fields to a 2D space by identifying key intersecting points in the DAG and calculating the scalar fields for every point, yielding a weighted Voronoi diagram like effect of space division. Two case studies demonstrate that the Topology Density Map supplies accurate information to users and provides an intuitive visualization for decision making. An interview with domain experts demonstrates the feasibility, usability, and effectiveness of our method.
Zezheng Feng, Haotian Li 0001, Wei Zeng 0004, Shuang-Hua Yang, Huamin Qu
IEEE Trans. Vis. Comput. Graph.3
2021 Exemplar-based Layout Fine-tuning for Node-link Diagrams
abstract
We design and evaluate a novel layout fine-tuning technique for node-link diagrams that facilitates exemplar-based adjustment of a group of substructures in batching mode. The key idea is to transfer user modifications on a local substructure to other substructures in the entire graph that are topologically similar to the exemplar. We first precompute a canonical representation for each substructure with node embedding techniques and then use it for on-the-fly substructure retrieval. We design and develop a light-weight interactive system to enable intuitive adjustment, modification transfer, and visual graph exploration. We also report some results of quantitative comparisons, three case studies, and a within-participant user study.
Jiacheng Pan, Wei Chen 0001, Shuyue Zhou, Wei Zeng 0004, Minfeng Zhu 0001, Jian Chen 0006, Siwei Fu, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.5
2021 Revisiting the Modifiable Areal Unit Problem in Deep Traffic Prediction with Visual Analytics
abstract
Deep learning methods are being increasingly used for urban traffic prediction where spatiotemporal traffic data is aggregated into sequentially organized matrices that are then fed into convolution-based residual neural networks. However, the widely known modifiable areal unit problem within such aggregation processes can lead to perturbations in the network inputs. This issue can significantly destabilize the feature embeddings and the predictions - rendering deep networks much less useful for the experts. This paper approaches this challenge by leveraging unit visualization techniques that enable the investigation of many-to-many relationships between dynamically varied multi-scalar aggregations of urban traffic data and neural network predictions. Through regular exchanges with a domain expert, we design and develop a visual analytics solution that integrates 1) a Bivariate Map equipped with an advanced bivariate colormap to simultaneously depict input traffic and prediction errors across space, 2) a Moran's I Scatterplot that provides local indicators of spatial association analysis, and 3) a Multi-scale Attribution View that arranges non-linear dot plots in a tree layout to promote model analysis and comparison across scales. We evaluate our approach through a series of case studies involving a real-world dataset of Shenzhen taxi trips, and through interviews with domain experts. We observe that geographical scale variations have important impact on prediction performances, and interactive visual exploration of dynamically varying inputs and outputs benefit experts in the development of deep traffic prediction models.
Wei Zeng 0004, Chengqiao Lin, Juncong Lin, Jincheng Jiang, Jiazhi Xia, Cagatay Turkay, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.1
2020 Visual Interpretation of Recurrent Neural Network on Multi-dimensional Time-series Forecast
abstract
Recent attempts at utilizing visual analytics to interpret Recurrent Neural Networks (RNNs) mainly focus on natural language processing (NLP) tasks that take symbolic sequences as input. However, many real-world problems like environment pollution forecasting apply RNNs on sequences of multi-dimensional data where each dimension represents an individual feature with semantic meaning such as PM2.5and SO2. RNN interpretation on multi-dimensional sequences is challenging as users need to analyze what features are important at different time steps to better understand model behavior and gain trust in prediction. This requires effective and scalable visualization methods to reveal the complex many-to-many relations between hidden units and features. In this work, we propose a visual analytics system to interpret RNNs on multi-dimensional time-series forecasts. Specifically, to provide an overview to reveal the model mechanism, we propose a technique to estimate the hidden unit response by measuring how different feature selections affect the hidden unit output distribution. We then cluster the hidden units and features based on the response embedding vectors. Finally, we propose a visual analytics system which allows users to visually explore the model behavior from the global and individual levels. We demonstrate the effectiveness of our approach with case studies using air pollutant forecast applications.
Qiaomu Shen, Yuzhe Jiang, Wei Zeng 0004, Alexis Kai-Hon Lau, Anna Vianova, Huamin Qu
PacificVis4
2020 SD-seq2seq : A Deep Learning Model for Bus Bunching Prediction Based on Smart Card Data
abstract
Bus bunching, a phenomenon due to the failure of headway or timetable adherence, often causes low level of public transit service with poor bus on-time performance and excessive passenger waiting time. To mitigate bus bunching, an accurate and real-time prediction method plays an important role. In this paper, we propose a supply-demand seq2seq model called SD-seq2seq to predict bus bunching using smart card data. Features from both supply and demand sides of bus service are taken into account, like bus stop type, dwelling time, passenger demand and type, and so on. Extensive experiments on multiple bus routes in real world demonstrate that our method outperforms other baseline methods. The proposed method is expected to provide useful online information of bus operation to both bus operators and passengers.
Zengyang Gong, Bo Du 0004, Zhidan Liu 0001, Wei Zeng 0004, Pascal Perez, Kaishun Wu
ICCCN4
2020 Deep Recognition of Vanishing-Point-Constrained Building Planes in Urban Street Views
abstract
This paper presents a new approach to recognizing vanishing-point-constrained building planes from a single image of street view. We first design a novel convolutional neural network (CNN) architecture that generates geometric segmentation of per-pixel orientations from a single street-view image. The network combines two-stream features of general visual cues and surface normals in gated convolution layers, and employs a deeply supervised loss that encapsulates multi-scale convolutional features. Our experiments on a new benchmark with fine-grained plane segmentations of real-world street views show that our network outperforms state-of-the-arts methods of both semantic and geometric segmentation. The pixel-wise segmentation exhibits coarse boundaries and discontinuities. We then propose to rectify the pixel-wise segmentation into perspectively-projected quads based on spatial proximity between the segmentation masks and exterior line segments detected through an image processing. We demonstrate how the results can be utilized to perspectively overlay images and icons on building planes in input photos, and provide visual cues for various applications.
Zhiliang Zeng, Mengyang Wu, Wei Zeng 0004, Chi-Wing Fu
IEEE Trans. Image Process.3
2020 LassoNet: Deep Lasso-Selection of 3D Point Clouds
abstract
Selection is a fundamental task in exploratory analysis and visualization of 3D point clouds. Prior researches on selection methods were developed mainly based on heuristics such as local point density, thus limiting their applicability in general data. Specific challenges root in the great variabilities implied by point clouds (e.g., dense vs. sparse), viewpoint (e.g., occluded vs. non-occluded), and lasso (e.g., small vs. large). In this work, we introduce LassoNet, a new deep neural network for lasso selection of 3D point clouds, attempting to learn a latent mapping from viewpoint and lasso to point cloud regions. To achieve this, we couple user-target points with viewpoint and lasso information through 3D coordinate transform and naive selection, and improve the method scalability via an intention filtering and farthest point sampling. A hierarchical network is trained using a dataset with over 30K lasso-selection records on two different point cloud data. We conduct a formal user study to compare LassoNet with two state-of-the-art lasso-selection methods. The evaluations confirm that our approach improves the selection effectiveness and efficiency across different combinations of 3D point clouds, viewpoints, and lasso selections. Project Website: https://LassoNet.github.io.
Chen Zhu-Tian, Wei Zeng 0004, Zhiguang Yang, Lingyun Yu 0005, Chi-Wing Fu, Huamin Qu
IEEE Trans. Vis. Comput. Graph.2
2019 VIStory: Interactive Storyboard for Exploring Visual Information in Scientific Publications
abstract
Many visual analytics have been developed for examining scientific publications comprising wealthy data such as authors and citations. The studies provide unprecedented insights on a variety of applications, e.g., literature review and collaboration analysis. However, visual information (i.e., figures) that are widely employed for storytelling and methods description are often neglected. We present VIStory, an interactive storyboard for exploring visual information in scientific publications. We harvest the data using an automatic figure extraction method, resulting in a large corpora of figures. Each figure contains various attributes such as dominant color and width/height ratio, together with faceted metadata of the publication including venues, authors, and keywords. To depict these information, we develop an intuitive interface consisting of three components: 1) Faceted View enables efficient query by publication metadata, benefiting from a nested table structure, 2) Storyboard View arranges paper rings -- a well-designed glyph for depicting figure attributes, in a themeriver layout to reveal temporal trends, and 3) Endgame View presents a highlighted figure together with the publication metadata. The system is especially useful for scientific publications containing substantial visual information, such as the visualization publications. We demonstrate the effectiveness of our approach using two case studies conducted on past ten-year IEEE VIS publications in 2009 - 2018.
Ao Dong, Wei Zeng 0004, Xi Chen 0072, Zhanglin Cheng
VINCI2
2019 Route-Aware Edge Bundling for Visualizing Origin-Destination Trails in Urban Traffic
abstract
Abstract Origin‐destination (OD) trails describe movements across space. Typical visualizations thereof use either straight lines or plot the actual trajectories. To reduce clutter inherent to visualizing large OD datasets, bundling methods can be used. Yet, bundling OD trails in urban traffic data remains challenging. Two specific reasons hereof are the constraints implied by the underlying road network and the difficulty of finding good bundling settings. To cope with these issues, we propose a new approach called Route Aware Edge Bundling (RAEB). To handle road constraints, we first generate a hierarchical model of the road‐and‐trajectory data. Next, we derive optimal bundling parameters, including kernel size and number of iterations, for a user‐selected level of detail of this model, thereby allowing users to explicitly trade off simplification vs accuracy. We demonstrate the added value of RAEB compared to state‐of‐the‐art trail bundling methods on both synthetic and real‐world traffic data for tasks that include the preservation of road network topology and the support of multiscale exploration.
Wei Zeng 0004, Qiaomu Shen, Yuzhe Jiang, Alexandru C. Telea
Comput. Graph. Forum1
2018 StreetVizor: Visual Exploration of Human-Scale Urban Forms Based on Street Views
abstract
Urban forms at human-scale, i.e., urban environments that individuals can sense (e.g., sight, smell, and touch) in their daily lives, can provide unprecedented insights on a variety of applications, such as urban planning and environment auditing. The analysis of urban forms can help planners develop high-quality urban spaces through evidence-based design. However, such analysis is complex because of the involvement of spatial, multi-scale (i.e., city, region, and street), and multivariate (e.g., greenery and sky ratios) natures of urban forms. In addition, current methods either lack quantitative measurements or are limited to a small area. The primary contribution of this work is the design of StreetVizor, an interactive visual analytics system that helps planners leverage their domain knowledge in exploring human-scale urban forms based on street view images. Our system presents two-stage visual exploration: 1) an AOI Explorer for the visual comparison of spatial distributions and quantitative measurements in two areas-of-interest (AOIs) at city- and region-scales; 2) and a Street Explorer with a novel parallel coordinate plot for the exploration of the fine-grained details of the urban forms at the street-scale. We integrate visualization techniques with machine learning models to facilitate the detection of street view patterns. We illustrate the applicability of our approach with case studies on the real-world datasets of four cities, i.e., Hong Kong, Singapore, Greater London and New York City. Interviews with domain experts demonstrate the effectiveness of our system in facilitating various analytical tasks.
Qiaomu Shen, Wei Zeng 0004, Yu Ye 0002, Stefan Müller Arisona, Simon Schubiger-Banz, Remo Aslak Burkhard, Huamin Qu
IEEE Trans. Vis. Comput. Graph.2
2017 Urban Fusion: Visualizing Urban Data Fused with Social Feeds via a Game Engine
abstract
This paper presents a framework which allows urban planners to navigate and interact with large datasets fused with social feeds in real-time, enhanced by a virtual reality (VR) capability, which further promotes the knowledge discovery process and allows to interact with urban data in natural yet immersive way. A challenge in urban planning is making decisions based on datasets which are many times ambiguous, together with effective use of newly available yet unstructured sources of information like social media. Providing expert users with novel ways of representing knowledge can be beneficial for decision making. Game engines have evolved into capable testbeds for novel visualization and interaction techniques. We therefore explore the possibility of using a modern game engine as a platform for knowledge representation in urban planning and how it can be used to model ambiguity. We also investigate how urban planners can benefit from immersion when it comes to data exploration and knowledge discovery. We apply the concept of using primitives to publicly available transportation datasets and social feeds of New York city, we discuss a gesture-based VR extension of our framework and lastly, we conclude the paper with feedback from expert users in urban planning and with an outlook of future challenges.
Ján Perhác, Wei Zeng 0004, Shiho Asada, Stefan Müller Arisona, Simon Schubiger-Banz, Remo Aslak Burkhard, Bernhard Klein
IV2
2017 Urban Mining: Visualizing the Availability of Construction Materials for Re-use in Future Cities
abstract
This paper presents a method and case study to visualize the urban stock of materials and its availability for use in building future cities. Re-using material from existing buildings for new buildings can be seen as a source for construction materials in times of depleting natural resources. The authors explain the concept of "urban mining" and the challenges, such as "How much resources are available in a city? Today? In the near future?" We explore what data are needed to answer the questions, and then discuss how to best visualize the data in an effective and intuitive way. We apply the concept to an exemplary real-world district in Singapore that is in transformation. Then, we discuss features of a visual tool prototype and explain the thinking behind the design, e.g., how the spatial and temporal dimensions can be presented. Lastly, we conclude the paper with an outlook of future challenges. The paper presents a multi-disciplinary approach with researchers from computer science, architecture, graphic design and material science, and contributes to the discussion of how to visualize knowledge and plan sustainable future cities.
Aurel Von Richthofen, Wei Zeng 0004, Shiho Asada, Remo Aslak Burkhard, Felix Heisel, Stefan Müller Arisona, Simon Schubiger-Banz
IV2
2017 Visualizing the Relationship Between Human Mobility and Points of Interest
abstract
In transportation studies, one fundamental problem is to analyze the departures and arrivals at locations in order to predict the travel demands for urban planning and traffic management. These movements can relate to many factors, e.g., activity distributions and household demographics. This paper presents how we use visualization to explore the relationship between people movements and activity distributions that are characterized by the points of interest (POIs). To effectively model and visualize such relationship, we introduce POI-mobility signature, a compact visual representation with two main components. 1) A mobility component to present major people movements information across temporal dimension. 2) A POI component to present the activity context over an area of interest in spatial domain. To derive the signature, we study assorted analytical tasks after discussing with transportation researchers, consider essential design principles, and apply the representation to study a real-world dataset, which is the massive public transportation data in Singapore with over 30 million trajectories and crowd-sourcing POIs retrieved from Foursquare. Finally, we conduct three case studies and interview three transportation experts to verify the efficacy of our method.
Wei Zeng 0004, Chi-Wing Fu, Stefan Müller Arisona, Simon Schubiger-Banz, Remo Aslak Burkhard, Kwan-Liu Ma
IEEE Trans. Intell. Transp. Syst.1
2017 A visual analytics design for studying rhythm patterns from human daily movement data
abstract
Human’s daily movements exhibit high regularity in a space–time context that typically forms circadian rhythms. Understanding the rhythms for human daily movements is of high interest to a variety of parties from urban planners, transportation analysts, to business strategists. In this paper, we present an interactive visual analytics design for understanding and utilizing data collected from tracking human’s movements. The resulting system identifies and visually presents frequent human movement rhythms to support interactive exploration and analysis of the data over space and time. Case studies using real-world human movement data, including massive urban public transportation data in Singapore and the MIT reality mining dataset, and interviews with transportation researches were conducted to demonstrate the effectiveness and usefulness of our system.
Wei Zeng 0004, Chi-Wing Fu, Stefan Müller Arisona, Simon Schubiger-Banz, Remo Aslak Burkhard, Kwan-Liu Ma
Vis. Informatics1
2016 Visualizing Waypoints-Constrained Origin-Destination Patterns for Massive Transportation Data
abstract
Abstract Origin‐destination (OD) pattern is a highly useful means for transportation research since it summarizes urban dynamics and human mobility. However, existing visual analytics are insufficient for certain OD analytical tasks needed in transport research. For example, transport researchers are interested in path‐related movements across congested roads, besides global patterns over the entire domain. Driven by this need, we proposewaypoints‐constrained OD visual analytics, a new approach for exploring path‐related OD patterns in an urban transportation network. First, we use hashing‐based query to support interactive filtering of trajectories through user‐specified waypoints. Second, we elaborate a set of design principles and rules, and derive a novel unified visual representation called thewaypoints‐constrained OD viewby carefully considering the OD flow presentation, the temporal variation, spatial layout and user interaction. Finally, we demonstrate the effectiveness of our interface with two case studies and expert interviews with five transportation experts.
Wei Zeng 0004, Chi-Wing Fu, Stefan Müller Arisona, Alexander Erath, Huamin Qu
Comput. Graph. Forum1
2014 Classifying watermelon ripeness by analysing acoustic signals using mobile devices
Wei Zeng 0004, Stefan Müller Arisona, Ian McLoughlin 0001
Pers. Ubiquitous Comput.1
2014 Visualizing Mobility of Public Transportation System
abstract
Public transportation systems (PTSs) play an important role in modern cities, providing shared/massive transportation services that are essential for the general public. However, due to their increasing complexity, designing effective methods to visualize and explore PTS is highly challenging. Most existing techniques employ network visualization methods and focus on showing the network topology across stops while ignoring various mobility-related factors such as riding time, transfer time, waiting time, and round-the-clock patterns. This work aims to visualize and explore passenger mobility in a PTS with a family of analytical tasks based on inputs from transportation researchers. After exploring different design alternatives, we come up with an integrated solution with three visualization modules: isochrone map view for geographical information, isotime flow map view for effective temporal information comparison and manipulation, and OD-pair journey view for detailed visual analysis of mobility factors along routes between specific origin-destination pairs. The isotime flow map linearizes a flow map into a parallel isoline representation, maximizing the visualization of mobility information along the horizontal time axis while presenting clear and smooth pathways from origin to destinations. Moreover, we devise several interactive visual query methods for users to easily explore the dynamics of PTS mobility over space and time. Lastly, we also construct a PTS mobility model from millions of real passenger trajectories, and evaluate our visualization techniques with assorted case studies with the transportation researchers.
Wei Zeng 0004, Chi-Wing Fu, Stefan Müller Arisona, Alexander Erath, Huamin Qu
IEEE Trans. Vis. Comput. Graph.1
2013 Visualizing Interchange Patterns in Massive Movement Data
abstract
Abstract Massive amount of movement data, such as daily trips made by millions of passengers in a city, are widely available nowadays. They are a highly valuable means not only for unveiling human mobility patterns, but also for assisting transportation planning, in particular for metropolises around the world. In this paper, we focus on a novel aspect of visualizing and analyzing massive movement data, i.e., the interchange pattern, aiming at revealing passenger redistribution in a traffic network. We first formulate a new model of circos figure, namely the interchange circos diagram, to present interchange patterns at a junction node in a bundled fashion, and optimize the color assignments to respect the connections within and between junction nodes. Based on this, we develop a family of visual analysis techniques to help users interactively study interchange patterns in a spatiotemporal manner: 1) multi‐spatial scales: from network junctions such as train stations to people flow across and between larger spatial areas; and 2) temporal changes of patterns from different times of the day. Our techniques have been applied to real movement data consisting of hundred thousands of trips, and we present also two case studies on how transportation experts worked with our interface.
Wei Zeng 0004, Chi-Wing Fu, Stefan Müller Arisona, Huamin Qu
Comput. Graph. Forum1