Bo Pan 0004

dblp:69/7781-4 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0009-4561-2469ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic Framework
abstract
Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep research and generate comprehensive reports. Despite its progress, existing deep research frameworks primarily focus on generating text-only content, leaving the automated generation of interleaved texts and visualizations underexplored. This novel task poses key challenges in designing informative visualizations and effectively integrating them with text reports. To address these challenges, we propose Formal Description of Visualization (FDV), a structured textual representation of charts that enables LLMs to learn from and generate diverse, high-quality visualizations. Building on this representation, we introduce Multimodal DeepResearcher, an agentic framework that decomposes the task into four stages: (1) researching, (2) exemplar report textualization, (3) planning and (4) multimodal report generation. For the evaluation of the generated reports, we develop MultimodalReportBench which contains 100 diverse topics as inputs, and a set of dedicated metrics for report and chart evaluation. Extensive experiments across models and evaluation methods demonstrate the effectiveness of Multimodal DeepResearcher. Notably, utilizing the same Claude 3.7 Sonnet model, Multimodal DeepResearcher achieves an 82% overall win rate over the baseline method.
Zhaorui Yang 0001, Bo Pan 0004, Yiyao Wang, Xingyu Liu 0003, Luoxuan Weng, Yingchaojie Feng, Haozhe Feng, Minfeng Zhu 0001, Wei Chen 0001
AAAI2
2026 IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
abstract
Yinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan, Yupeng Xie, Jiale Lao, Yiyao Wang, Haoxuan Li, Tingting Gao, Bo Pan, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu, Yingchaojie Feng, Yuyu Luo, Wei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yinghao Tang, Xueding Liu, Tingfeng Lan, Jiale Lao, Yiyao Wang, Tingting Gao, Bo Pan 0004, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu 0001, Yingchaojie Feng, Yuyu Luo, Wei Chen 0001
ACL (1)10
2026 DreamJourney: Perpetual View Generation With Video Diffusion Models
abstract
Perpetual view generation aims to synthesize a long-term video corresponding to an arbitrary camera trajectory solely from a single input image. Recent methods commonly utilize a pre-trained text-to-image diffusion model to synthesize new content of previously unseen regions along camera movement. However, the underlying 2D diffusion model lacks 3D awareness and results in distorted artifacts. Moreover, they are limited to generating views of static 3D scenes, neglecting to capture object movements within the dynamic 4D world. To alleviate these issues, we present DreamJourney, a two-stage framework that leverages the world simulation capacity of video diffusion models to trigger a new perpetual scene view generation task with both camera movements and object dynamics. Specifically, in stage I, DreamJourney first lifts the input image to 3D point cloud and renders a sequence of partial images from a specific camera trajectory. A video diffusion model is then utilized as generative prior to complete the missing regions and enhance visual coherence across the sequence, producing a cross-view consistent video adheres to the 3D scene and camera trajectory. Meanwhile, we introduce two simple yet effective strategies (early stopping and view padding) to further stabilize the generation process and improve visual quality. Next, in stage II, DreamJourney leverages a multimodal large language model to produce a text prompt describing object movements in current view, and uses video diffusion model to animate current view with object movements. Stage I and II are repeated recurrently, enabling perpetual dynamic scene view generation. Extensive experiments demonstrate the superiority of our DreamJourney over state-of-the-art methods both quantitatively and qualitatively. Our project page:https://dream-journey.vercel.app/.
Bo Pan 0004, Yang Chen 0048, Yingwei Pan, Ting Yao 0003, Wei Chen 0001, Tao Mei 0001
IEEE Trans. Multim.1
2026 VidGuard3D: A Visual Risk Analysis Approach for Protecting 3D Assets Against Video-Based Reconstruction Attacks
abstract
The unauthorized acquisition of 3D assets by means of advanced techniques in 3D reconstruction is a major but often overlooked threat to publishers of videos. Preventing such threats is challenging due to the uninterpretable nature of 3D reconstruction and the diversity in requirements of 3D model demonstration. In this paper, we introduce VidGuard3D-a visual risk analysis approach that quantifies and locates the sources of risk for video-based 3D asset reconstruction attacks. Our approach uses attack simulation to support users in formulating a comprehensive understanding of 3D asset leakage risks, with a particular focus on the correlations between video segments and exposure risks of user-specified areas in the asset. We also proposed a prototype system that integrates this approach to facilitate video editing according to the knowledge of correlations. Two operations of video editing can be swiftly applied by users to form editing plans and minimize detected leakage risks. Finally, we conducted a user study and case studies that demonstrated the practicality and effectiveness of our approach.
Yiyao Wang, Ollie Woodman, Shenghui Hu, Ruizhe Pan, Bo Pan 0004, Xumeng Wang, Minfeng Zhu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2026 Exploring Multimodal Prompt for Visualization Authoring With Large Language Models
abstract
Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in precision and expressiveness for conveying visualization intent, leading to misinterpretation and time-consuming iterations. To address these limitations, we conduct an empirical study to understand how LLMs interpret ambiguous or incomplete text prompts in the context of visualization authoring, and the conditions making LLMs misinterpret user intent. Informed by the findings, we introduce visual prompts as a complementary input modality to text prompts, which help clarify user intent and improve LLMs' interpretation abilities. To explore the potential of multimodal prompting in visualization authoring, we design VisPilot, which enables users to easily create visualizations using multimodal prompts, including text, sketches, and direct manipulations on existing visualizations. We evaluate VisPilot through a controlled user study and an expert evaluation. The results suggest that multimodal prompts facilitate users in communicating spatial constraints, local references, and design preferences while maintaining comparable task efficiency to text-only prompting. We further discuss when text, visual, and hybrid prompts are beneficial for visualization authoring, and summarize design implications for future human-AI authoring systems.
Zhen Wen 0001, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Bo Pan 0004, Minfeng Zhu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2025 XGraphRAG: Interactive Visual Analysis for Graph-based Retrieval-Augmented Generation
abstract
Graph-based Retrieval-Augmented Generation (RAG) has shown great capability in enhancing Large Language Model (LLM)’s answer with an external knowledge base. Compared to traditional RAG, it introduces a graph as an intermediate representation to capture better structured relational knowledge in the corpus, elevating the precision and comprehensiveness of generation results. However, developers usually face challenges in analyzing the effectiveness of GraphRAG on their dataset due to GraphRAG’s complex information processing pipeline and the overwhelming amount of LLM invocations involved during graph construction and query, which limits GraphRAG interpretability and accessibility. This research proposes a visual analysis framework that helps RAG developers identify critical recalls of GraphRAG and trace these recalls through the GraphRAG pipeline. Based on this framework, we develop XGraphRAG, a prototype system incorporating a set of interactive visualizations to facilitate users’ analysis process, boosting failure cases collection and improvement opportunities identification. Our evaluation demonstrates the effectiveness and usability of our approach. Our work is open-sourced and available at https://github.com/Gk0Wk/XGraphRAG.
Bo Pan 0004, Yingchaojie Feng, Jieyi Chen, Minfeng Zhu 0001, Wei Chen 0001
PacificVis2
2025 X-SARF: Semantically-Aware Neural Radiance Fields for Sparse-View X-Ray Novel View Synthesis
Zhide Jiang, Wenan Zhang, Bin Geng, Yayi Xia, Bo Pan 0004, Wei Chen 0001
PRCV (14)7
2025 AgentCoord: Visually exploring coordination strategy for LLM-based multi-agent collaboration
Bo Pan 0004, Jiaying Lu 0005, Zhen Wen 0001, Yingchaojie Feng, Minfeng Zhu 0001, Wei Chen 0001
Comput. Graph.1
2025 AgentLens: Visual Analysis for Agent Behaviors in LLM-Based Autonomous Systems
abstract
Recently, Large Language Model based Autonomous System (LLMAS) has gained great popularity for its potential to simulate complicated behaviors of human societies. One of its main challenges is to present and analyze the dynamic events evolution of LLMAS. In this work, we present a visualization approach to explore the detailed statuses and agents' behavior within LLMAS. Our approach outlines a general pipeline that organizes raw execution events from LLMAS into a structured behavior model. We leverage a behavior summarization algorithm to create a hierarchical summary of these behaviors, arranged according to their sequence over time. Additionally, we design a cause trace method to mine the causal relationship between agent behaviors. We then develop AgentLens, a visual analysis system that leverages a hierarchical temporal visualization for illustrating the evolution of LLMAS, and supports users to interactively investigate details and causes of agents' behaviors. Two usage scenarios and a user study demonstrate the effectiveness and usability of our AgentLens.
Jiaying Lu 0005, Bo Pan 0004, Jieyi Chen, Yingchaojie Feng, Yuchen Peng, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.2
2024 XNLI: Explaining and Diagnosing NLI-Based Visual Data Analysis
abstract
Natural language interfaces (NLIs) enable users to flexibly specify analytical intentions in data visualization. However, diagnosing the visualization results without understanding the underlying generation process is challenging. Our research explores how to provide explanations for NLIs to help users locate the problems and further revise the queries. We present XNLI, an explainable NLI system for visual data analysis. The system introduces a Provenance Generator to reveal the detailed process of visual transformations, a suite of interactive widgets to support error adjustments, and a Hint Generator to provide query revision hints based on the analysis of user queries and interactions. Two usage scenarios of XNLI and a user study verify the effectiveness and usability of the system. Results suggest that XNLI can significantly enhance task accuracy without interrupting the NLI-based analysis process.
Yingchaojie Feng, Xingbo Wang 0001, Bo Pan 0004, Kamkwai Wong, Yuxin Ma 0001, Huamin Qu, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.3
2024 Differentiable Design Galleries: A Differentiable Approach to Explore the Design Space of Transfer Functions
abstract
The transfer function is crucial for direct volume rendering (DVR) to create an informative visual representation of volumetric data. However, manually adjusting the transfer function to achieve the desired DVR result can be time-consuming and unintuitive. In this paper, we propose Differentiable Design Galleries, an image-based transfer function design approach to help users explore the design space of transfer functions by taking advantage of the recent advances in deep learning and differentiable rendering. Specifically, we leverage neural rendering to learn a latent design space, which is a continuous manifold representing various types of implicit transfer functions. We further provide a set of interactive tools to support intuitive query, navigation, and modification to obtain the target design, which is represented as a neural-rendered design exemplar. The explicit transfer function can be reconstructed from the target design with a differentiable direct volume renderer. Experimental results on real volumetric data demonstrate the effectiveness of our method.
Bo Pan 0004, Jiaying Lu 0005, Weifeng Chen 0003, Yiyao Wang, Minfeng Zhu 0001, Chenhao Yu, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.1