VLDB 2026 Research / reviewers in the wild / expert
Sicheng Song
dblp:265/3568
· DBLP profile ↗
12ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-2158-0353ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VizQStudio: Iterative Visualization Literacy MCQs Design With Simulated StudentsabstractMultiple-choice questions (MCQs) are a widely used educational tool, particularly in domains such as visualization literacy that require broad conceptual coverage and support diverse real-world applications. However, designing high-quality visualization literacy MCQs remains challenging, as instructors must coordinate multimodal elements (e.g., charts, question stems, and distractors), address diverse visualization tasks, and accommodate learners with heterogeneous backgrounds. Existing visualization literacy assessments primarily rely on standardized, fixed item banks, offering limited support for iterative question design that adapts to differences in learners' abilities, backgrounds, and reasoning strategies. To address these challenges, we present VizQStudio, a visual analytics system that supports instructors in iteratively designing and refining visualization literacy MCQs using MLLM-powered simulated students. Instructors can specify diverse student profiles spanning demographics, knowledge levels, and learning-related traits. The system then visualizes how simulated students reason about and respond to different question components, helping instructors explore potential misconceptions, difficulty calibration, and design trade-offs prior to classroom deployment. We investigate VizQStudio through a mixed-method evaluation, including expert interviews, case studies, a classroom deployment, and a large-scale online study. Our results indicate that MCQs designed with VizQStudio can support measurable learning gains and, within our exploratory online sample, yielded observed post-test outcomes similar to established benchmark questions, while enabling greater flexibility and scalability during the design process. Overall, this work reframes MLLM-based student simulation in assessment authoring as a design-time, exploratory aid. By examining both its value and limitations in realistic instructional settings, we surface design insights that inform how future systems can support instructor-centered, iterative, and responsible uses of AI for multimodal assessment design in visualization literacy and related domains. Zixin Chen, Yuhang Zeng, Sicheng Song, Yanna Lin, Huamin Qu, Meng Xia 0002 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | VizDefender: Unmasking Visualization Tampering Through Proactive Localization and Intent InferenceabstractThe integrity of data visualizations is increasingly threatened by image editing techniques that enable subtle yet deceptive tampering. Through a formative study, we define this challenge and categorize tampering techniques into two primary types: data manipulation and visual encoding manipulation. To address this, we present VizDefender, a framework for tampering detection and analysis. The framework integrates two core components: 1) a semi-fragile watermark module that protects the visualization by embedding a location map to images, which allows for the precise localization of tampered regions while preserving visual quality, and 2) an intent analysis module that leverages Multimodal Large Language Models (MLLMs) to interpret manipulation, inferring the attacker's intent and misleading effects. Extensive evaluations and user studies demonstrate the effectiveness of our methods. Sicheng Song, Zixin Chen, Huamin Qu, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question AnsweringabstractMisleading visualizations, which manipulate chart representations to support specific claims, can distort perception and lead to incorrect conclusions.Despite decades of research, they remain a widespread issue, posing risks to public understanding and raising safety concerns for AI systems involved in data-driven communication.While recent multimodal large language models (MLLMs) show strong chart comprehension abilities, their capacity to detect and interpret misleading charts remains unexplored.We introduce Misleading ChartQA benchmark, a large-scale multimodal dataset designed to evaluate MLLMs on misleading chart reasoning.It contains 3,026 curated examples spanning 21 misleader types and 10 chart types, each with standardized chart code, CSV data, multiple-choice questions, and labeled explanations, validated through iterative MLLM checks and expert human review.We benchmark 24 state-of-the-art MLLMs, analyze their performance across misleader types and chart formats, and propose a novel regionaware reasoning pipeline that enhances model accuracy.Our work lays the foundation for developing MLLMs that are robust, trustworthy, and aligned with the demands of responsible visual communication. Zixin Chen, Sicheng Song, KaShun Shum, Yanna Lin, Rui Sheng, Huamin Qu |
EMNLP | 2 |
| 2025 | RhythmTA: A Visual-Aided Interactive System for ESL Rhythm Training via Dubbing Practice
Chang Chen 0005, Sicheng Song, Shuchang Xu, Huamin Qu, Yanna Lin |
UIST | 2 |
| 2025 | GVVST: Image-Driven Style Extraction From Graph Visualizations for Visual Style TransferabstractIncorporating automatic style extraction and transfer from existing well-designed graph visualizations can significantly alleviate the designer's workload. There are many types of graph visualizations. In this paper, our work focuses on node-link diagrams. We present a novel approach to streamline the design process of graph visualizations by automatically extracting visual styles from well-designed examples and applying them to other graphs. Our formative study identifies the key styles that designers consider when crafting visualizations, categorizing them into global and local styles. Leveraging deep learning techniques such as saliency detection models and multi-label classification models, we develop end-to-end pipelines for extracting both global and local styles. Global styles focus on aspects such as color scheme and layout, while local styles are concerned with the finer details of node and edge representations. Through a user study and evaluation experiment, we demonstrate the efficacy and time-saving benefits of our method, highlighting its potential to enhance the graph visualization design process. Sicheng Song, Yanna Lin, Huamin Qu, Changbo Wang, Chenhui Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | GraphDecoder: Recovering Diverse Network Graphs From Visualization Images via Attention-Aware LearningabstractDNGs are diverse network graphs with texts and different styles of nodes and edges, including mind maps, modeling graphs, and flowcharts. They are high-level visualizations that are easy for humans to understand but difficult for machines. Inspired by the process of human perception of graphs, we propose a method called GraphDecoder to extract data from raster images. Given a raster image, we extract the content based on a neural network. We built a semantic segmentation network based on U-Net. We increase the attention mechanism module, simplify the network model, and design a specific loss function to improve the model's ability to extract graph data. After this semantic segmentation network, we can extract the data of all nodes and edges. We then combine these data to obtain the topological relationship of the entire DNG. We also provide an interactive interface for users to redesign the DNGs. We verify the effectiveness of our method by evaluations and user studies on datasets collected on the internet and generated datasets. Sicheng Song, Chenhui Li 0001, Juntong Chen, Changbo Wang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | GVQA: Learning to Answer Questions about Graphs with Visualizations via Knowledge BaseabstractGraphs are common charts used to represent the topological relationship between nodes. It is a powerful tool for data analysis and information retrieval tasks involve asking questions about graphs. In formative study, we found that questions for graphs are not only about the relationship of nodes but also about the properties of graph elements. We propose a pipeline to answer natural language questions about graph visualizations and generate visual answers. We first extract the data from graphs and convert them into GML format. We design data structures to encode graph information and convert them into an knowledge base. We then extract topic entities from questions. We feed questions, entities and knowledge bases into our question-answer model to obtain the SPARQL queries for textual answers. Finally, we design a module to present the answers visually. A user study demonstrates that these visual and textual answers are useful, credible and and transparent. Sicheng Song, Juntong Chen, Chenhui Li 0001, Changbo Wang |
CHI | 1 |
| 2023 | iARVis: Mobile AR Based Declarative Information Visualization Authoring, Exploring and SharingabstractWe present iARVis, a proof-of-concept toolkit for creating, experiencing, and sharing mobile AR-based information visualization environments. Over the past years, AR has emerged as a promising medium for information and data visualization beyond the physical media and the desktop, enabling interactivity and eliminating spatial limits. However, the creation of such environments remains difficult and frequently necessitates low-level programming expertise and lengthy hand encodings. We present a declarative approach for defining the augmented reality (AR) environment, including how information is automatically positioned, laid out, and interacted with, to improve the efficiency and flexibility of constructing AR-based information visualization environments. We provide fundamental layout and visual components such as the grid, rich text, images, and charts for the development of complex visualization widgets, as well as automatic targeting methods based on image and object tracking for the development of the AR environment. To increase design efficiency, we also provide features such as hot-reload and several creation levels for both novice and advanced users. We also investigate how the augmented reality-based visualization environment could persist and be shared through the internet and provide ways for storing, sharing, and restoring the environment to give a continuous and seamless experience. To demonstrate the viability and extensibility, we evaluate iARVis using a variety of use cases along with performance evaluation and expert reviews. Chenhui Li 0001, Sicheng Song, Changbo Wang |
VR | 3 |
| 2023 | VividGraph: Learning to Extract and Redesign Network Graphs From Visualization ImagesabstractNetwork graphs are common visualization charts. They often appear in the form of bitmaps in articles, web pages, magazine prints, and designer sketches. People often want to modify graphs because of their poor design, but it is difficult to obtain their underlying data. In this article, we present VividGraph, a pipeline for automatically extracting and redesigning graphs from static images. We propose using convolutional neural networks to solve the problem of graph data extraction. Our method is robust to hand-drawn graphs, blurred graph images, and large graph images. We also present a graph classification module to make it effective for directed graphs. We propose two evaluation methods to demonstrate the effectiveness of our approach. It can be used to quickly transform designer sketches, extract underlying data from existing graphs, and interactively redesign poorly designed graphs. Sicheng Song, Chenhui Li 0001, Yujing Sun 0003, Changbo Wang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | OpinionManager: Visual Exploration of Online Reviews in P2P AccommodationabstractUser-generated online reviews are critical in P2P accommodations. They contain a wealth of information about the opinions and experiences of users, which help better understand consumer decisions and improve products and services. However, the huge volume of reviews makes it difficult for potential customers to gain useful insights and for managers to track customer opinions. To address these problems, we first use topic modeling techniques for customer opinion mining. Then, we build a deep learning network for sentiment analysis. Finally, we perform sentiment analysis of the reviews at the aspect level to obtain the sentiment vector representation of the accommodation. Moreover, we design a visual analytic system with a user-friendly interface to facilitate interactive analysis. Evaluation including user and case studies demonstrates the usefulness and effectiveness of this system. Changbo Wang, Sicheng Song, Kirlin Li, Chenhui Li 0001 |
VINCI | 4 |
| 2021 | Effect of Antenna Orientation on the Air-to-Air Channel in Arbitrary 3D SpaceabstractUnmanned Aerial Vehicles (UAVs) often lack the size, weight, and power to support large antenna arrays or a large number of radio chains. Despite such limitations, emerging applications that require the use of swarms, where UAVs form a pattern and coordinate towards a common goal, must have the capability to transmit in any direction in three-dimensional (3D) space from moment to moment. In this work, we design a measurement study to evaluate the role of antenna polarization diversity on UAV systems communicating in arbitrary 3D space. To do so, we construct flight patterns where one transmitting UAV is hovering at a high altitude (80 m) and a receiving UAV hovers at 114 different positions that span 3D space at a radial distance of approximately 20 m along equally-spaced elevation and azimuth angles. To understand the role of diverse antenna polarizations, both UAVs have a horizontally-mounted antenna and a vertically-mounted antenna-each attached to a dedicated radio chain-creating four wireless channels. With this measurement campaign, we seek to understand how to optimally select an antenna orientation and quantify the gains in such selections. N. Cameron Matson, Syed Muhammad Hashir, Sicheng Song, Dinesh Rajan, Joseph David Camp |
WOWMOM | 3 |
| 2020 | Interactive Attention Model Explorer for Natural Language Processing Tasks with Unbalanced Data SizesabstractConventional attention visualization tools compromise either the readability or the information conveyed when documents are lengthy, especially when these documents have imbalanced sizes. Our work strives toward a more intuitive visualization for a subset of Natural Language Processing tasks, where attention is mapped between documents with imbalanced sizes. We extend the flow map visualization to enhance the readability of the attention-augmented documents. Through interaction, our design enables semantic filtering that helps users prioritize important tokens and meaningful matching for an in-depth exploration. Case studies and informal user studies in machine comprehension prove that our visualization effectively helps users gain initial understandings about what their models are "paying attention to." We discuss how the work can be extended to other domains, as well as being plugged into more end-to-end systems for model error analysis. Zhihang Dong, Sherry Tongshuang Wu, Sicheng Song |
PacificVis | 3 |