VLDB 2026 Research / reviewers in the wild / expert
Zekai Shao 0001
dblp:342/9205-1
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-2014-5293ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From łog π to π: Taming Divergence in Soft Clipping via Bilateral Decoupled Decay of Probability Gradient WeightabstractXiaoliang Fu, Jiaye Lin, Yangyi Fang, Chaowen Hu, Cong Qin, Zekai Shao, Binbin Zheng, Lu Pan, Ke Zeng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoliang Fu, Jiaye Lin, Yangyi Fang, Chaowen Hu, Cong Qin, Zekai Shao 0001 |
ACL (1) | 6 |
| 2026 | MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM ReasoningabstractXiaoliang Fu, Jiaye Lin, Yangyi Fang, Binbin Zheng, Chaowen Hu, Zekai Shao, Cong Qin, Lu Pan, Ke Zeng, Xunliang Cai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoliang Fu, Jiaye Lin, Yangyi Fang, Chaowen Hu, Zekai Shao 0001, Cong Qin |
ACL (1) | 6 |
| 2026 | NotebookRAG: Retrieving Multiple Notebooks to Augment the Generation of EDA Notebooks for Crowd-WisdomabstractHigh-quality exploratory data analysis (EDA) is essential in the data science pipeline, but remains highly dependent on analysts' expertise and effort. While recent LLM-based approaches partially reduce this burden, they struggle to generate effective analysis plans and appropriate insights and visualizations when user intent is abstract. Meanwhile, a vast collection of analysis notebooks produced across platforms and organizations contains rich analytical knowledge that can potentially guide automated EDA. Retrieval-augmented generation (RAG) provides a natural way to leverage such corpora, but general methods often treat notebooks as static documents and fail to fully exploit their potentially knowledge for automating EDA. To address these limitations, we propose NotebookRAG, a method that takes user intent, datasets, and existing notebooks as input to retrieve, enhance, and reuse relevant notebook content for automated EDA generation. For retrieval, we transform code cells into context-enriched executable components, which improve retrieval quality and enable rerun with new data to generate updated visualizations and reliable insights. For generation, an agent leverages enhanced retrieval content to construct effective EDA plans, derive insights, and produce appropriate visualizations. Evidence from a user study with 24 participants confirms the superiority of our method in producing high-quality and intent-aligned EDA notebooks. Yi Shan 0002, Zekai Shao 0001, Kai Xu 0003, Siming Chen 0001 |
PacificVis | 3 |
| 2025 | Unlocking Scientific Concepts: How Effective Are LLM-Generated Analogies for Student Understanding and Classroom Practice?
Zekai Shao 0001, Deqing Yang, Siming Chen 0001 |
CHI | 1 |
| 2025 | Fine-Tuned Large Language Model for Visualization System: A Study on Self-Regulated Learning in EducationabstractLarge Language Models (LLMs) have shown great potential in intelligent visualization systems, especially for domain-specific applications. Integrating LLMs into visualization systems presents challenges, and we categorize these challenges into three alignments: domain problems with LLMs, visualization with LLMs, and interaction with LLMs. To achieve these alignments, we propose a framework and outline a workflow to guide the application of fine-tuned LLMs to enhance visual interactions for domain-specific tasks. These alignment challenges are critical in education because of the need for an intelligent visualization system to support beginners' self-regulated learning. Therefore, we apply the framework to education and introduce Tailor-Mind, an interactive visualization system designed to facilitate self-regulated learning for artificial intelligence beginners. Drawing on insights from a preliminary study, we identify self-regulated learning tasks and fine-tuning objectives to guide visualization design and tuning data construction. Our focus on aligning visualization with fine-tuned LLM makes Tailor-Mind more like a personalized tutor. Tailor-Mind also supports interactive recommendations to help beginners better achieve their learning goals. Model performance evaluations and user studies confirm that Tailor-Mind improves the self-regulated learning experience, effectively validating the proposed framework. Zekai Shao 0001, Ziyue Lin, Shengbin Yue, Chiokit Leong, Rory James Zauner, Zhongyu Wei, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Narrative Player: Reviving Data Narratives With VisualsabstractData-rich documents are commonly found across various fields such as business, finance, and science. However, a general limitation of these documents for reading is their reliance on text to convey data and facts. Visual representation of text aids in providing a satisfactory reading experience in comprehension and engagement. However, existing work emphasizes presenting the insights within phrases or sentences, rather than fully conveying data stories within the whole paragraphs and engaging readers. To provide readers with satisfactory data stories, this paper presents Narrative Player, a novel method that automatically revives data narratives with consistent and contextualized visuals. Specifically, it accepts a paragraph and corresponding data table as input and leverages LLMs to characterize the clauses and extract contextualized data facts. Subsequently, the facts are transformed into a coherent visualization sequence with a carefully designed optimization-based approach. Animations are also assigned between adjacent visualizations to enable seamless transitions. Finally, the visualization sequence, transition animations, and audio narration generated by text-to-speech technologies are rendered into a data video. The evaluation results showed that the automatic-generated data videos were well-received by participants and experts for enhancing reading. Zekai Shao 0001, Leixian Shen, Haotian Li 0001, Yi Shan 0002, Huamin Qu, Yun Wang 0012, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | ChartInsighter: An Approach for Mitigating Hallucination in Time-Series Chart Summary Generation With a Benchmark DatasetabstractEffective chart summary can significantly reduce the time and effort decision makers spend interpreting charts, enabling precise and efficient communication of data insights. Previous studies have faced challenges in generating accurate and semantically rich summaries of time-series data charts. In this paper, we identify summary elements and common hallucination types in the generation of time-series chart summaries, which serve as our guidelines for automatic generation. We introduce ChartInsighter, which automatically generates chart summaries of time-series data, effectively reducing hallucinations in chart summary generation. Specifically, we assign multiple agents to generate the initial chart summary and collaborate iteratively, during which they invoke external data analysis modules to extract insights and compile them into a coherent summary. Additionally, we implement a self-consistency test method to validate and correct our summary. We create a high-quality benchmark of charts and summaries, with hallucination types annotated on a sentence-by-sentence basis, facilitating the evaluation of the effectiveness of reducing hallucinations. Our evaluations using our benchmark show that our method surpasses state-of-the-art models, and that our summary hallucination rate is the lowest, which effectively reduces various hallucinations and improves summary quality. Bomiao Wang, Xueli Shu, Zhen Liu 0058, Zekai Shao 0001, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | LEVA: Using Large Language Models to Enhance Visual AnalyticsabstractVisual analytics supports data analysis tasks within complex domain problems. However, due to the richness of data types, visual designs, and interaction designs, users need to recall and process a significant amount of information when they visually analyze data. These challenges emphasize the need for more intelligent visual analytics methods. Large language models have demonstrated the ability to interpret various forms of textual data, offering the potential to facilitate intelligent support for visual analytics. We propose LEVA, a framework that uses large language models to enhance users' VA workflows at multiple stages: onboarding, exploration, and summarization. To support onboarding, we use large language models to interpret visualization designs and view relationships based on system specifications. For exploration, we use large language models to recommend insights based on the analysis of system status and data to facilitate mixed-initiative exploration. For summarization, we present a selective reporting strategy to retrace analysis history through a stream visualization and generate insight reports with the help of large language models. We demonstrate how LEVA can be integrated into existing visual analytics systems. Two usage scenarios and a user study suggest that LEVA effectively aids users in conducting visual analytics. Yuheng Zhao, Yu Zhang 0043, Zekai Shao 0001, Cagatay Turkay, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Interpreting Autonomous Driving Corner Cases: A Visual Analytics ApproachabstractWith the progression of artificial intelligence, there has been substantial advancement in autonomous driving technology. However, even the most advanced systems may confront failures in certain corner cases, necessitating enhanced analytical approaches. Traditional approaches focused on the numerical analysis of isolated sensor data, are often insufficient for deriving meaningful insights in such situations. To address this inadequacy, we propose a visual analytics approach, crafted to aid domain experts in performing analyses and extracting system improvements from cases with unexpected behaviors. This approach intricately integrates extensive driving scenarios and low-level module behaviors into the autonomous driving decision-making process, utilizing rich visualizations and an interface for interactive exploration and systematic synthesis of findings. Uniquely, our system opens the "black box" of modules in the decision-making pipeline during corner cases, taking into account both the overall decision-making pipeline and the fine-grained behaviors of the modules in the pipeline, setting our approach apart from previous works. To validate our system’s effectiveness, we perform two case studies, inviting domain experts for evaluation, and the results confirm our system’s efficacy in allowing experts to obtain crucial insights into autonomous driving systems. Zekai Shao 0001, Xingyu Qiu, Linbing Xiang, Siming Chen 0001 |
PacificVis | 2 |
| 2024 | TransforLearn: Interactive Visual Tutorial for the Transformer ModelabstractThe widespread adoption of Transformers in deep learning, serving as the core framework for numerous large-scale language models, has sparked significant interest in understanding their underlying mechanisms. However, beginners face difficulties in comprehending and learning Transformers due to its complex structure and abstract data representation. We present TransforLearn, the first interactive visual tutorial designed for deep learning beginners and non-experts to comprehensively learn about Transformers. TransforLearn supports interactions for architecture-driven exploration and task-driven exploration, providing insight into different levels of model details and their working processes. It accommodates interactive views of each layer's operation and mathematical formula, helping users to understand the data flow of long text sequences. By altering the current decoder-based recursive prediction results and combining the downstream task abstractions, users can deeply explore model processes. Our user study revealed that the interactions of TransforLearn are positively received. We observe that TransforLearn facilitates users' accomplishment of study tasks and a grasp of key concepts in Transformer effectively. Zekai Shao 0001, Ziqin Luo, Haibo Hu 0002, Cagatay Turkay, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Visual Explanation for Open-Domain Question Answering With BERTabstractOpen-domain question answering (OpenQA) is an essential but challenging task in natural language processing that aims to answer questions in natural language formats on the basis of large-scale unstructured passages. Recent research has taken the performance of benchmark datasets to new heights, especially when these datasets are combined with techniques for machine reading comprehension based on Transformer models. However, as identified through our ongoing collaboration with domain experts and our review of literature, three key challenges limit their further improvement: (i) complex data with multiple long texts, (ii) complex model architecture with multiple modules, and (iii) semantically complex decision process. In this paper, we present VEQA, a visual analytics system that helps experts understand the decision reasons of OpenQA and provides insights into model improvement. The system summarizes the data flow within and between modules in the OpenQA model as the decision process takes place at the summary, instance and candidate levels. Specifically, it guides users through a summary visualization of dataset and module response to explore individual instances with a ranking visualization that incorporates context. Furthermore, VEQA supports fine-grained exploration of the decision flow within a single module through a comparative tree visualization. We demonstrate the effectiveness of VEQA in promoting interpretability and providing insights into model enhancement through a case study and expert evaluation. Zekai Shao 0001, Shuran Sun, Yuheng Zhao, Siyuan Wang 0025, Zhongyu Wei, Tao Gui, Cagatay Turkay, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |