Shishi Xiao

dblp:342/9310 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2025
0009-0008-0262-5289ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation
abstract
Recently, state-of-the-art text-to-image generation models, such as FLUX and Ideogram 2.0, have made significant progress in sentence-level visual text rendering. In this paper, we focus on the more challenging scenarios of article-level visual text rendering and address a novel task of generating high-quality business content, including infographics and slides, based on user provided article-level descriptive prompts and ultra-dense layouts. The fundamental challenges are twofold: significantly longer context lengths and the scarcity of high-quality business content data.In contrast to most previous works that focus on a limited number of sub-regions and sentence-level prompts, ensuring precise adherence to ultra-dense layouts with tens or even hundreds of sub-regions in business content is far more challenging. We make two key technical contributions: (i) the construction of scalable, high-quality business content dataset, i.e., Infographics-650K, equipped with ultra-dense layouts and prompts by implementing a layer-wise retrieval-augmented infographic generation scheme; and (ii) a layout-guided cross attention scheme, which injects tens of region-wise prompts into a set of cropped region latent space according to the ultra-dense layouts, and refine each sub-regions flexibly during inference using a layout conditional CFG. We demonstrate the strong results of our system compared to previous SOTA systems such as FLUX and SD3 on our BizEval prompt set. Additionally, we conduct thorough ablation experiments to verify the effectiveness of each component. We hope our constructed Infographics-650K and BizEval can encourage the broader community to advance the progress of business content generation.
Yuyang Peng, Shishi Xiao, Keming Wu, Qisheng Liao, Danqing Huang, Yuhui Yuan
CVPR2
2025 VizTA: Enhancing Comprehension of Distributional Visualization with Visual-Lexical Fused Conversational Interface
abstract
Abstract Comprehending visualizations requires readers to interpret visual encoding and the underlying meanings actively. This poses challenges for visualization novices, particularly when interpreting distributional visualizations that depict statistical uncertainty. Advancements in LLM‐based conversational interfaces show promise in promoting visualization comprehension. However, they fail to provide contextual explanations at fine‐grained granularity, and chart readers are still required to mentally bridge visual information and textual explanations during conversations. Our formative study highlights the expectations for both lexical and visual feedback, as well as the importance of explicitly linking these two modalities throughout the conversation. The findings motivate the design of VizTA, a visualization teaching assistant that leverages the fusion of visual and lexical feedback to help readers better comprehend visualization. VizTA features a semantic‐aware conversational agent capable of explaining contextual information within visualizations and employs a visual‐lexical fusion design to facilitate chart‐centered conversation. A between‐subject study with 24 participants demonstrates the effectiveness of VizTA in supporting the understanding and reasoning tasks of distributional visualization across multiple scenarios.
Liangwei Wang 0001, Zhan Wang 0001, Shishi Xiao, Le Liu 0008, Fugee Tsung, Wei Zeng 0004
Comput. Graph. Forum3
2025 ModalChorus: Visual Probing and Alignment of Multi-Modal Embeddings via Modal Fusion Map
abstract
Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting in decreased model performance and diminished generalization. To address this problem, we design ModalChorus, an interactive system for visual probing and alignment of multi-modal embeddings. ModalChorus primarily offers a two-stage process: 1) embedding probing with Modal Fusion Map (MFM), a novel parametric dimensionality reduction method that integrates both metric and nonmetric objectives to enhance modality fusion; and 2) embedding alignment that allows users to interactively articulate intentions for both point-set and set-set alignments. Quantitative and qualitative comparisons for CLIP embeddings with existing dimensionality reduction (e.g., t-SNE and MDS) and data fusion (e.g., data context map) methods demonstrate the advantages of MFM in showcasing cross-modal features over common vision-language datasets. Case studies reveal that ModalChorus can facilitate intuitive discovery of misalignment and efficient re-alignment in scenarios ranging from zero-shot classification to cross-modal retrieval and generation.
Shishi Xiao, Xingchen Zeng, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.2
2024 TypeDance: Creating Semantic Typographic Logos from Image through Personalized Generation
abstract
Semantic typographic logos harmoniously blend typeface and imagery to represent semantic concepts while maintaining legibility. Conventional methods using spatial composition and shape substitution are hindered by the conflicting requirement for achieving seamless spatial fusion between geometrically dissimilar typefaces and semantics. While recent advances made AI generation of semantic typography possible, the end-to-end approaches exclude designer involvement and disregard personalized design. This paper presents TypeDance, an AI-assisted tool incorporating design rationales with the generative model for personalized semantic typographic logo design. It leverages combinable design priors extracted from uploaded image exemplars and supports type-imagery mapping at various structural granularity, achieving diverse aesthetic designs with flexible control. Additionally, we instantiate a comprehensive design workflow in TypeDance, including ideation, selection, generation, evaluation, and iteration. A two-task user evaluation, including imitation and creation, confirmed the usability of TypeDance in design across different usage scenarios.
Shishi Xiao, Liangwei Wang 0001, Xiaojuan Ma, Wei Zeng 0004
CHI1
2024 The Contemporary Art of Image Search: Iterative User Intent Expansion via Vision-Language Model
abstract
Image search is an essential and user-friendly method to explore vast galleries of digital images. However, existing image search methods heavily rely on proximity measurements like tag matching or image similarity, requiring precise user inputs for satisfactory results. To meet the growing demand for a contemporary image search engine that enables accurate comprehension of users' search intentions, we introduce an innovative user intent expansion framework. Our framework leverages visual-language models to parse and compose multi-modal user inputs to provide more accurate and satisfying results. It comprises two-stage processes: 1) a parsing stage that incorporates a language parsing module with large language models to enhance the comprehension of textual inputs, along with a visual parsing module that integrates an interactive segmentation module to swiftly identify detailed visual elements within images; and 2) a logic composition stage that combines multiple user search intents into a unified logic expression for more sophisticated operations in complex searching scenarios. Moreover, the intent expansion framework enables users to perform flexible contextualized interactions with the search results to further specify or adjust their detailed search intents iteratively. We implemented the framework into an image search system for NFT (non-fungible token) search and conducted a user study to evaluate its usability and novel properties. The results indicate that the proposed framework significantly improves users' image search experience. Particularly the parsing and contextualized interactions prove useful in allowing users to express their search intents more accurately and engage in a more enjoyable iterative search experience.
Qian Zhu 0010, Shishi Xiao, Kang Zhang 0001, Wei Zeng 0004
Proc. ACM Hum. Comput. Interact.3
2024 MetroBUX: A Topology-Based Visual Analytics for Bus Operational Uncertainty EXploration
abstract
In the public transportation system, punctuality benefits both bus operation and passengers’ travel experience. However, uncertainty exists due to complex traffic conditions and heterogeneous driving behaviors. To analyze bus operational uncertainty, transport planners and bus operators need a tool that supports multi-granular modeling, spatio-temporal representation, and interactive exploration. To meet the requirement, we present MetroBUX, a visual analytics system for$B$us operational$U$ncertainty e$X$ploration. MetroBUX aligns daily bus trips and models stop-level uncertainty of bus arrival time. It has a consolidated interface with three main views: Map View for presenting the spatial distribution of uncertainty, Temporal View for tracking the evolution of uncertainty, and Trip View for inspecting uncertainty propagation. Specifically, MetroBUX enables integrated spatio-temporal analysis by connecting topological uncertainty distribution at different periods in a nested tracking graph. Furthermore, it supports interactive and hierarchical exploration, including region-, route-, trip-, and stop-level analysis. Case studies on real-world bus operational data and domain experts’ feedback demonstrate the efficiency of MetroBUX.
Shishi Xiao, Lingdan Shao, Bo Du 0004, Yang Wang 0006, Qiaomu Shen, Wei Zeng 0004
IEEE Trans. Intell. Transp. Syst.1
2024 Let the Chart Spark: Embedding Semantic Context into Chart with Text-to-Image Generative Model
abstract
Pictorial visualization seamlessly integrates data and semantic context into visual representation, conveying complex information in an engaging and informative manner. Extensive studies have been devoted to developing authoring tools to simplify the creation of pictorial visualizations. However, mainstream works follow a retrieving-and-editing pipeline that heavily relies on retrieved visual elements from a dedicated corpus, which often compromise data integrity. Text-guided generation methods are emerging, but may have limited applicability due to their predefined entities. In this work, we propose ChartSpark, a novel system that embeds semantic context into chart based on text-to-image generative models. ChartSpark generates pictorial visualizations conditioned on both semantic context conveyed in textual inputs and data information embedded in plain charts. The method is generic for both foreground and background pictorial generation, satisfying the design practices identified from empirical research into existing pictorial visualizations. We further develop an interactive visual interface that integrates a text analyzer, editing module, and evaluation module to enable users to generate, modify, and assess pictorial visualizations. We experimentally demonstrate the usability of our tool, and conclude with a discussion of the potential of using text-to-image generative models combined with an interactive interface for visualization design.
Shishi Xiao, Suizi Huang, Wei Zeng 0004
IEEE Trans. Vis. Comput. Graph.1
2024 Generative AI for visualization: State of the art and future directions
abstract
Generative AI (GenAI) has witnessed remarkable progress in recent years and demonstrated impressive performance in various generation tasks in different domains such as computer vision and computational design. Many researchers have attempted to integrate GenAI into visualization framework, leveraging the superior generative capacity for different operations. Concurrently, recent major breakthroughs in GenAI like diffusion model and large language model have also drastically increase the potential of GenAI4VIS. From a technical perspective, this paper looks back on previous visualization studies leveraging GenAI and discusses the challenges and opportunities for future research. Specifically, we cover the applications of different types of GenAI methods including sequence, tabular, spatial and graph generation techniques for different tasks of visualization which we summarize into four major stages: data enhancement, visual mapping generation, stylization and interaction. For each specific visualization sub-task, we illustrate the typical data and concrete GenAI algorithms, aiming to provide in-depth understanding of the state-of-the-art GenAI4VIS techniques and their limitations. Furthermore, based on the survey, we discuss three major aspects of challenges and research opportunities including evaluation, dataset, and the gap between end-to-end GenAI methods and visualizations. By summarizing different generation algorithms, their current applications and limitations, this paper endeavors to provide useful insights for future GenAI4VIS research.
Jianing Hao, Yihan Hou, Zhan Wang 0001, Shishi Xiao, Yuyu Luo, Wei Zeng 0004
Vis. Informatics5
2023 CP-NeRF: Conditionally Parameterized Neural Radiance Fields for Cross-scene Novel View Synthesis
abstract
Abstract Neural radiance fields (NeRF) have demonstrated a promising research direction for novel view synthesis. However, the existing approaches either require per‐scene optimization that takes significant computation time or condition on local features which overlook the global context of images. To tackle this shortcoming, we propose the Conditionally Parameterized Neural Radiance Fields (CP‐NeRF), a plug‐in module that enables NeRF to leverage contextual information from different scales. Instead of optimizing the model parameters of NeRFs directly, we train a Feature Pyramid hyperNetwork (FPN) that extracts view‐dependent global and local information from images within or across scenes to produce the model parameters. Our model can be trained end‐to‐end with standard photometric loss from NeRF. Extensive experiments demonstrate that our method can significantly boost the performance of NeRF, achieving state‐of‐the‐art results in various benchmark datasets.
Hao He 0011, Yixun Liang, Shishi Xiao, Jierun Chen, Ying-Cong Chen
Comput. Graph. Forum3
2023 WYTIWYR: A User Intent-Aware Framework with Multi-modal Inputs for Visualization Retrieval
abstract
Abstract Retrieving charts from a large corpus is a fundamental task that can benefit numerous applications such as visualization recommendations. The retrieved results are expected to conform to both explicit visual attributes (e.g., chart type, colormap) and implicit user intents (e.g., design style, context information) that vary upon application scenarios. However, existing example‐based chart retrieval methods are built upon non‐decoupled and low‐level visual features that are hard to interpret, while definition‐based ones are constrained to pre‐defined attributes that are hard to extend. In this work, we propose a new framework, namelyWYTIWYR (What‐You‐Think‐Is‐What‐You‐Retrieve), that integrates user intents into the chart retrieval process. The framework consists of two stages: first, theAnnotationstage disentangles the visual attributes within the query chart; and second, theRetrievalstage embeds the user's intent with customized text prompt as well as bitmap query chart, to recall targeted retrieval result. We develop aprototypeWYTIWYRsystem leveraging a contrastive language‐image pre‐training (CLIP) model to achieve zero‐shot classification as well as multi‐modal input encoding, and test the prototype on a large corpus with charts crawled from the Internet. Quantitative experiments, case studies, and qualitative interviews are conducted. The results demonstrate the usability and effectiveness of our proposed framework.
Shishi Xiao, Yihan Hou, Cheng Jin 0003, Wei Zeng 0004
Comput. Graph. Forum1
2021 Progressive Band-Separated Convolutional Neural Network for Multispectral Pansharpening
abstract
Recently, convolutional neural networks (CNNs) have been introduced to pansharpening for enhancing fusion accuracy and overcoming the drawbacks of the conventional methods. However, most of methods based on CNN fail to distinguish the difference of multispectral bands, and only use a uniform set of convolutional kernels to extract features. In this paper, we design a progressive, band-separated convolutional network architecture for discriminatively learning the features and relation among spectral bands, aiming to address the problem mentioned before. More specifically, the proposed architecture mainly consists of three aspects. First, to accurately preserve the spectral peculiarities, we divide the multispectral input image in terms of its bands into several groups. Second, our original panchromatic and multispectral inputs are filtered by a high-pass operation to further yield more spatial details. Third, we use a spectral fusion module (SFM) for each group and associate them to progressively assemble the whole architecture. It is worth mentioning that the architecture could be integrated into any other competitive CNNs to improve the performance. Both visual and quantitative experiments have demonstrated that our proposed method outperforms recent state-of-the-art pansharpening techniques.
Shishi Xiao, Cheng Jin 0003, Tianjing Zhang, Ran Ran 0001, Liang-Jian Deng
IGARSS1