Minfeng Zhu 0001

dblp:163/6031-1 · DBLP profile ↗
← Back
50ranked-venue papers
5as first author
43since 2021 · last 2026
0000-0002-6711-3099ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 3 first-author · 30 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic Framework
abstract
Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep research and generate comprehensive reports. Despite its progress, existing deep research frameworks primarily focus on generating text-only content, leaving the automated generation of interleaved texts and visualizations underexplored. This novel task poses key challenges in designing informative visualizations and effectively integrating them with text reports. To address these challenges, we propose Formal Description of Visualization (FDV), a structured textual representation of charts that enables LLMs to learn from and generate diverse, high-quality visualizations. Building on this representation, we introduce Multimodal DeepResearcher, an agentic framework that decomposes the task into four stages: (1) researching, (2) exemplar report textualization, (3) planning and (4) multimodal report generation. For the evaluation of the generated reports, we develop MultimodalReportBench which contains 100 diverse topics as inputs, and a set of dedicated metrics for report and chart evaluation. Extensive experiments across models and evaluation methods demonstrate the effectiveness of Multimodal DeepResearcher. Notably, utilizing the same Claude 3.7 Sonnet model, Multimodal DeepResearcher achieves an 82% overall win rate over the baseline method.
Zhaorui Yang 0001, Bo Pan 0004, Yiyao Wang, Xingyu Liu 0003, Luoxuan Weng, Yingchaojie Feng, Haozhe Feng, Minfeng Zhu 0001, Wei Chen 0001
AAAI9
2026 IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation
abstract
Yinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan, Yupeng Xie, Jiale Lao, Yiyao Wang, Haoxuan Li, Tingting Gao, Bo Pan, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu, Yingchaojie Feng, Yuyu Luo, Wei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yinghao Tang, Xueding Liu, Tingfeng Lan, Jiale Lao, Yiyao Wang, Tingting Gao, Bo Pan 0004, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu 0001, Yingchaojie Feng, Yuyu Luo, Wei Chen 0001
ACL (1)13
2026 Attention as Selector: Unlocking VLM Attention for Long Document Page Retrieval
abstract
Visual Language Models (VLMs) have become a robust foundation for document question answering.Processing long documents remains challenging due to limited context windows and computational budgets.Existing page-level retrieval methods offer a practical solution, typically encoding pages and queries into vectors and ranking them via cosine similarity.However, such embedding-based methods (i) lack query-page interaction before similarity scoring and (ii) usually require a large-scale dataset to align visual and textual embeddings.In this paper, we observe that the cross-modal attention maps of well-trained VLMs are able to highlight semantically relevant regions.Building on this insight, we present CAPS (Crossmodal Attention as Page Selector), a retrieval framework that utilizes attention mechanisms inside VLMs for page selection.Specifically, CAPS first enhances attention-based retrieval capability with a small amount of contrastive data, then identifies the most effective attention head through expert head selection, and finally employs an adaptive filtering mechanism to obtain an appropriate number of relevant page candidates.Extensive experiments on four long-document benchmarks demonstrate that CAPS outperforms state-of-the-art embedding-based methods in both retrieval precision and downstream DocQA accuracy.Notably, CAPS achieves these gains using less than 10% of the training data required by competing baselines, highlighting the data efficiency of attention-based page retrieval.
Minfeng Zhu 0001, Linxin Bao, Wei Chen 0001, Linchao Zhu
ACL (1)1
2026 ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models
abstract
Large language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks. Understanding how LLMs internally represent knowledge remains a significant challenge. Despite Sparse Autoencoders (SAEs) have emerged as a promising technique for extracting interpretable features from LLMs, SAE features do not inherently align with human-understandable concepts, making their interpretation cumbersome and labor-intensive. To bridge the gap between SAE features and human concepts, we present ConceptViz, a visual analytics system designed for exploring concepts in LLMs. ConceptViz implements a novel Identification ⇒ Interpretation ⇒Validation pipeline, enabling users to query SAEs using concepts of interest, interactively explore concept-to-feature alignments, and validate the correspondences through model behavior verification. We demonstrate the effectiveness of ConceptViz through two usage scenarios and a user study. Our results show that ConceptViz enhances interpretability research by streamlining the discovery and validation of meaningful concept representations in LLMs, ultimately aiding researchers in building more accurate mental models of LLM features. Our code and user guide are publicly available at https://github.com/Happy-Hippo209/conceptViz.
Zhen Wen 0001, Qiqi Jiang, Chenxiao Li, Yiyao Wang, Xiuqi Huang, Minfeng Zhu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.9
2026 RAGExplorer: A Visual Analytics System for the Comparative Diagnosis of RAG Systems
abstract
The advent of Retrieval-Augmented Generation (RAG) has significantly enhanced the ability of Large Language Models (LLMs) to produce factually accurate and up-to-date responses. However, the performance of a RAG system is not determined by a single component but emerges from a complex interplay of modular choices, such as embedding models and retrieval algorithms. This creates a vast and often opaque configuration space, making it challenging for developers to understand performance trade-offs and identify optimal designs. To address this challenge, we present RAGExplorer, a visual analytics system for the systematic comparison and diagnosis of RAG configurations. RAGExplorer guides users through a seamless macro-to-micro analytical workflow. Initially, it empowers developers to survey the performance landscape across numerous configurations, allowing for a high-level understanding of which design choices are most effective. For a deeper analysis, the system enables users to drill down into individual failure cases, investigate how differences in retrieved information contribute to errors, and interactively test hypotheses by manipulating the provided context to observe the resulting impact on the generated answer. We demonstrate the effectiveness of RAGExplorer through detailed case studies and user studies, validating its ability to empower developers in navigating the complex RAG design space. Our code and user guide are publicly available at https://github.com/Thymezzz/RAGExplorer.
Yingchaojie Feng, Zhen Wen 0001, Minfeng Zhu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.5
2026 VidGuard3D: A Visual Risk Analysis Approach for Protecting 3D Assets Against Video-Based Reconstruction Attacks
abstract
The unauthorized acquisition of 3D assets by means of advanced techniques in 3D reconstruction is a major but often overlooked threat to publishers of videos. Preventing such threats is challenging due to the uninterpretable nature of 3D reconstruction and the diversity in requirements of 3D model demonstration. In this paper, we introduce VidGuard3D-a visual risk analysis approach that quantifies and locates the sources of risk for video-based 3D asset reconstruction attacks. Our approach uses attack simulation to support users in formulating a comprehensive understanding of 3D asset leakage risks, with a particular focus on the correlations between video segments and exposure risks of user-specified areas in the asset. We also proposed a prototype system that integrates this approach to facilitate video editing according to the knowledge of correlations. Two operations of video editing can be swiftly applied by users to form editing plans and minimize detected leakage risks. Finally, we conducted a user study and case studies that demonstrated the practicality and effectiveness of our approach.
Yiyao Wang, Ollie Woodman, Shenghui Hu, Ruizhe Pan, Bo Pan 0004, Xumeng Wang, Minfeng Zhu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.8
2026 QuRAFT: Enhancing Quantum Algorithm Design by Visual Linking Between Mathematical Concepts and Quantum Circuits
abstract
The emergence of quantum computers heralds a new frontier in computational power, empowering quantum algorithms to address challenges that defy classical computation. However, the design of quantum algorithms is challenging as it largely requires the manual efforts of quantum experts to transit mathematical expressions to quantum circuit diagrams. To ease this process, particularly for prototyping, educational, and modular design workflows, we propose to bridge the textual and visual contexts between mathematics and quantum circuits through visual linking and transitions. We contribute a design space for quantum algorithm design, focusing on the textual and visual elements, interactions, and design patterns throughout the quantum algorithm design process. Informed by the design space, we introduce QuRAFT, a visual interface that facilitates a seamless transition from abstract mathematical expressions to concrete quantum circuits. QuRAFT incorporates a suite of eight integrated visual and interaction designs tailored to support users in the formulation, implementation, and validation process of the quantum algorithm design. Through two detailed case studies and a user evaluation, this paper demonstrates the effectiveness of QuRAFT. Feedback from quantum computing experts highlights the practical utility of QuRAFT in algorithm design and provides valuable implications for future advancements in visualization and interaction design within the quantum computing domain.
Zhen Wen 0001, Jieyi Chen, Siwei Tan, Jianwei Yin, Minfeng Zhu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2026 Exploring Multimodal Prompt for Visualization Authoring With Large Language Models
abstract
Recent advances in large language models (LLMs) have shown great potential in automating the process of visualization authoring through simple natural language utterances. However, instructing LLMs using natural language is limited in precision and expressiveness for conveying visualization intent, leading to misinterpretation and time-consuming iterations. To address these limitations, we conduct an empirical study to understand how LLMs interpret ambiguous or incomplete text prompts in the context of visualization authoring, and the conditions making LLMs misinterpret user intent. Informed by the findings, we introduce visual prompts as a complementary input modality to text prompts, which help clarify user intent and improve LLMs' interpretation abilities. To explore the potential of multimodal prompting in visualization authoring, we design VisPilot, which enables users to easily create visualizations using multimodal prompts, including text, sketches, and direct manipulations on existing visualizations. We evaluate VisPilot through a controlled user study and an expert evaluation. The results suggest that multimodal prompts facilitate users in communicating spatial constraints, local references, and design preferences while maintaining comparable task efficiency to text-only prompting. We further discuss when text, visual, and hybrid prompts are beneficial for visualization authoring, and summarize design implications for future human-AI authoring systems.
Zhen Wen 0001, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Bo Pan 0004, Minfeng Zhu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.7
2026 A Utility-Aware Privacy-Preserving Method for Trajectory Publication
abstract
The security of individual privacy is paramount for trajectory publication, while preserving trajectory utility is also essential to serve analysis tasks such as urban planning and transportation development. To assess and maintain trajectory utility, existing studies consider geographic context. However, innate semantic characteristics of trajectories (e.g., origin, destination, stay point, path) have been overlooked, which prevents data owners from specifying task-specific utility measurements and, consequently, from achieving a delicate balance between privacy and utility. This paper proposes an interactive trajectory publishing approach driven by flexible utility considerations, which processes trajectory points according to their semantics to fulfill diverse utility requirements. Concretely, we decouple trajectories into origin-destination (OD) and path components: ODs are generalized into regions to satisfy $k$k-anonymity, and paths are sanitized within each OD group using road-network-aware differential privacy under predefined privacy constraints. We also develop a visual interface to support exploration and comprehension of privacy-preserving solutions, through which we incorporate human knowledge into the privacy scheme. Experiments on real-world urban datasets demonstrate the effectiveness of our approach.
Ziliang Wu, Xumeng Wang, Zhaosong Huang, Tiansheng Zhang, Minfeng Zhu 0001, Xiuqi Huang, Mingliang Xu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.5
2026 G2: A customizable web-based framework for authoring interactive visualizations
abstract
Specifying visual encodings and interactions is exhausting but essential for authoring interactive visualizations. In this paper, we present G2, a customizable framework designed to support rapid generation of interactive visualizations with uniform specifications. G2 employs a data-driven grammar of graphics and defines a uniform set of interaction specifications. We discuss the design and implementation of G2 with rich examples, and a user interview to demonstrate its effectiveness. Since its first release in March 2016, G2 has undergone 360 iterations, received 12,000 stars, and supported over 22,600 related projects on Github.
Bairui Su, Zhifeng Lin, Xiaojuan Liao, Zihan Zhou 0009, Minfeng Zhu 0001, Wei Chen 0001
Vis. Informatics6
2026 Spatially-anchored visual diagnostics for smart construction
Erqing Zhang, Longting Zhu, Minfeng Zhu 0001, Wei Chen 0001
Vis. Informatics3
2025 Don't Reinvent the Wheel: Efficient Instruction-Following Text Embedding based on Guided Space Transformation
abstract
Yingchaojie Feng, Yiqun Sun, Yandong Sun, Minfeng Zhu, Qiang Huang, Anthony Kum Hoe Tung, Wei Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yingchaojie Feng, Yiqun Sun, Yandong Sun, Minfeng Zhu 0001, Anthony K. H. Tung, Wei Chen 0001
ACL (1)4
2025 R1-Onevision: Advancing Generalized Multimodal Reasoning Through Cross-Modal Formalization
Yi Yang 0001, Xiaoxuan He, Hongkun Pan, Xiyan Jiang, Xingtao Yang, Haoyu Lu, Dacheng Yin, Fengyun Rao, Minfeng Zhu 0001, Wei Chen 0001
ICCV10
2025 DataLab: A Unified Platform for LLM-Powered Business Intelligence
abstract
Business intelligence (BI) transforms large volumes of data within modern organizations into actionable insights for informed decision-making. Recently, large language model (LLM)-based agents have streamlined the BI workflow by automatically performing task planning, reasoning, and actions in executable environments based on natural language (NL) queries. However, existing approaches primarily focus on individual BI tasks such as NL2SQL and NL2VIS. The fragmentation of tasks across different data roles and tools lead to inefficiencies and potential errors due to the iterative and collaborative nature of BI. In this paper, we introduce DataLab, a unified BI platform that integrates a one-stop LLM-based agent framework with an augmented computational notebook interface. DataLab supports various BI tasks for different data roles in data preparation, analysis, and visualization by seamlessly combining LLM assistance with user customization within a single environment. To achieve this unification, we design a domain knowledge incorporation module tailored for enterprise-specific BI tasks, an inter-agent communication mechanism to facilitate information sharing across the BI workflow, and a cell-based context management strategy to enhance context utilization efficiency in BI notebooks. Extensive experiments demonstrate that DataLab achieves state-of-the-art performance on various BI tasks across popular research benchmarks. Moreover, DataLab maintains high effectiveness and efficiency on real-world datasets from Tencent, achieving up to a 58.58% increase in accuracy and a 61.65 % reduction in token cost on enterprise-specific BI tasks.
Luoxuan Weng, Yinghao Tang, Yingchaojie Feng, Zhuo Chang, Ruiqin Chen, Haozhe Feng, Chen Hou, Danqing Huang, Yang Li 0106, Huaming Rao, Canshi Wei, Xiuqi Huang, Minfeng Zhu 0001, Yuxin Ma 0001, Bin Cui 0001, Peng Chen 0021, Wei Chen 0001
ICDE17
2025 MOTION: Multi-object Video Editing with Training-Free Attention Guidance
Qitong Yan, Jian Jia, Bo Wang 0071, Quan Chen 0006, Peng Jiang 0002, Minfeng Zhu 0001, Linchao Zhu, Wei Chen 0001
ICIC (18)8
2025 RG3: Mitigating Memorization of Graph Diffusion Model in One Denoising Step
Yijing Liu 0003, Minfeng Zhu 0001, Wei Chen 0001
ICIC (10)4
2025 CultiVerse: Towards Cross-Cultural Understanding for Paintings with Large Language Model
abstract
Understanding cultural heritage through technology faces challenges in connecting with diverse audiences, especially when interpreting art across cultures. In this work, we present CultiVerse, a visual analytics system that leverages Large Language Models (LLMs) to support cross-cultural appreciation of Traditional Chinese Paintings (TCPs). CultiVerse operates within a mixed-initiative framework and guides users through three stages: extracting cultural context, aligning cross-cultural symbols, and extrapolating meaning in the viewer's cultural frame. By combining an interactive interface with LLM-powered analysis, the system enables deeper engagement with symbolic meanings and encourages serendipitous cross-cultural discoveries. Our approach bridges AI interpretation and human insight to foster mutual understanding in a multicultural setting. A curated TCP dataset supports exploration, while empirical evaluations confirm that CultiVerse enhances user understanding, interpretation accuracy, and cultural empathy.
Wei Zhang 0219, Kamkwai Wong, Biying Xu, Yiwen Ren, Yuhuai Li, Yingchaojie Feng, Minfeng Zhu 0001, Wei Chen 0001
ACM Multimedia7
2025 XGraphRAG: Interactive Visual Analysis for Graph-based Retrieval-Augmented Generation
abstract
Graph-based Retrieval-Augmented Generation (RAG) has shown great capability in enhancing Large Language Model (LLM)’s answer with an external knowledge base. Compared to traditional RAG, it introduces a graph as an intermediate representation to capture better structured relational knowledge in the corpus, elevating the precision and comprehensiveness of generation results. However, developers usually face challenges in analyzing the effectiveness of GraphRAG on their dataset due to GraphRAG’s complex information processing pipeline and the overwhelming amount of LLM invocations involved during graph construction and query, which limits GraphRAG interpretability and accessibility. This research proposes a visual analysis framework that helps RAG developers identify critical recalls of GraphRAG and trace these recalls through the GraphRAG pipeline. Based on this framework, we develop XGraphRAG, a prototype system incorporating a set of interactive visualizations to facilitate users’ analysis process, boosting failure cases collection and improvement opportunities identification. Our evaluation demonstrates the effectiveness and usability of our approach. Our work is open-sourced and available at https://github.com/Gk0Wk/XGraphRAG.
Bo Pan 0004, Yingchaojie Feng, Jieyi Chen, Minfeng Zhu 0001, Wei Chen 0001
PacificVis6
2025 CausalPrism: A visual analytics approach for subgroup-based causal heterogeneity exploration
Xingyu Liu 0003, Jiehui Zhou, Xumeng Wang, Kamkwai Wong, Wei Zhang 0219, Juntian Zhang, Minfeng Zhu 0001, Wei Chen 0001
Comput. Graph.7
2025 AgentCoord: Visually exploring coordination strategy for LLM-based multi-agent collaboration
Bo Pan 0004, Jiaying Lu 0005, Zhen Wen 0001, Yingchaojie Feng, Minfeng Zhu 0001, Wei Chen 0001
Comput. Graph.7
2025 Visual analysis approach for mutual fund selection
Fan Yan, Yong Wang 0021, Xuanwu Yue, Kamkwai Wong, Ketian Mao, Rong Zhang 0011, Huamin Qu, Minfeng Zhu 0001, Wei Chen 0001
Frontiers Comput. Sci.9
2025 A Summarization-Based Pattern-Aware Matrix Reordering Approach
Zihan Zhou 0009, Jiacheng Pan, Xumeng Wang, Dongming Han, Fangzhou Guo, Minfeng Zhu 0001, Wei Chen 0001
J. Comput. Sci. Technol.6
2025 JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
abstract
The proliferation of large language models (LLMs) has underscored concerns regarding their security vulnerabilities, notably against jailbreak attacks, where adversaries design jailbreak prompts to circumvent safety mechanisms for potential misuse. Addressing these concerns necessitates a comprehensive analysis of jailbreak prompts to evaluate LLMs' defensive capabilities and identify potential weaknesses. However, the complexity of evaluating jailbreak performance and understanding prompt characteristics makes this analysis laborious. We collaborate with domain experts to characterize problems and propose an LLM-assisted framework to streamline the analysis process. It provides automatic jailbreak assessment to facilitate performance evaluation and support analysis of components and keywords in prompts. Based on the framework, we design JailbreakLens, a visual analysis system that enables users to explore the jailbreak performance against the target model, conduct multi-level analysis of prompt characteristics, and refine prompt instances to verify findings. Through a case study, technical evaluations, and expert interviews, we demonstrate our system's effectiveness in helping users evaluate model security and identify model weaknesses.
Yingchaojie Feng, Zhizhang (David) Chen, Zhining Kang, Wei Zhang 0219, Minfeng Zhu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.7
2025 Graphon-Based Visual Abstraction for Large Multi-Layer Networks
abstract
Graph visualization techniques provide a foundational framework for offering comprehensive overviews and insights into cloud computing systems, facilitating efficient management and ensuring their availability and reliability. Despite the enhanced computational and storage capabilities of larger-scale cloud computing architectures, they introduce significant challenges to traditional graph-based visualization due to issues of hierarchical heterogeneity, scalability, and data incompleteness. This paper proposes a novel abstraction approach to visualize large multi-layer networks. Our method leverages graphons, a probabilistic representation of network layers, to encompass three core steps: an inner-layer summary to identify stable and volatile substructures, an inter-layer mixup for aligning heterogeneous network layers, and a context-aware multi-layer joint sampling technique aimed at reducing network scale while retaining essential topological characteristics. By abstracting complex network data into manageable weighted graphs, with each graph depicting a distinct network layer, our approach renders these intricate systems accessible on standard computing hardware. We validate our methodology through case studies, quantitative experiments and expert evaluations, demonstrating its effectiveness in managing large multi-layer networks, as well as its applicability to broader network types such as transportation and social networks.
Ziliang Wu, Minfeng Zhu 0001, Zhaosong Huang, Junxu Chen, Tiansheng Zhang, Shengbing Shi, Qiang Bai, Hongchao Qu, Xiuqi Huang, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.2
2025 A human-centric perspective on interpretability in large language models
Zihan Zhou 0009, Minfeng Zhu 0001, Wei Chen 0001
Vis. Informatics2
2025 STEP-LINK: STEP-by-Step Tutorial Editing with Programmable LINKages
abstract
Programming tutorials serve a crucial role in teaching coding and programming techniques. Creating high-quality programming tutorials remains a laborious task. Authors devote effort in writing step-by-step solutions, creating examples, and editing existing tutorials. We explore the potential of using the text-code connection to improve the authoring experience of programming tutorials. We proposed a mixed-initiative approach to infer, establish, and maintain the latent text-code connections. With a series of interactions, the STEP-LINK ( STEP- by-Step Tutorial Editing with Programmable LINK ages ) prototype leverages text-code connections to assist users in authoring tutorials. The results of our experiment demonstrate the effectiveness of our system in supporting users in the authoring of step-by-step code explanations, the creation of examples, and the iteration of tutorials.
Junming Ke, Zhen Wen 0001, Junhua Lu, Biao Zhu, Minfeng Zhu 0001, Wei Chen 0001
Vis. Informatics7
2025 FundSelector: A visual analysis system for mutual fund selection
abstract
Mutual funds are one of the most important and popular investment ways for ordinary investors to maintain and increase the value of their assets. However, it is challenging for ordinary investors to select optimal mutual funds from thousands of fund choices managed by different managers. Various investors often have different personal investment preferences and it is difficult to characterize their preferences quickly. Also, mutual fund performance relies on various factors (e.g., the economic market and the management of fund managers), and most of these factors are dynamically changing, making it difficult to efficiently compare different mutual funds in detail. To address these challenges, we propose FundSelector, an interactive multi-view visual analytics system that quantifies user preferences to rank mutual funds and allows ordinary investors to explore mutual fund performance in terms of multiple factors and scales. Two novel visual designs are proposed to enable detailed comparisons of mutual funds. Rank-informed bipartite contribution bar chart provides interpretable fund ranking results by explicitly showing both positive and negative factors. Elastic trend chart allows investors to analyze and compare the temporal evolution of the mutual funds’ performances in a customizable way. We evaluated FundSelector through two case studies and interviews with eight ordinary investors. The results highlight its effectiveness and utility.
Fan Yan, Yong Wang 0021, Xuanwu Yue, Kamkwai Wong, Ketian Mao, Rong Zhang 0011, Huamin Qu, Minfeng Zhu 0001, Wei Chen 0001
Vis. Informatics9
2024 Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning
abstract
Zhaorui Yang, Tianyu Pang, Haozhe Feng, Han Wang, Wei Chen, Minfeng Zhu, Qian Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhaorui Yang 0001, Tianyu Pang, Haozhe Feng, Wei Chen 0001, Minfeng Zhu 0001, Qian Liu 0033
ACL (1)6
2024 Nuwa: An Authoring Tool for Graph Visualizations
abstract
Authoring graph visualization requires advanced programming skills, expert domain knowledge, and significant workload. Existing authoring tools either support limited templates of graph visualization, or suffer from a high learning cost. We analyze the design requirements on a tool for graph visualizations, and contribute Nuwa, a user-friendly declarative authoring tool for the interactive specification of graph visualizations in terms of data, entity, change, and encoding. Our implementation empowers users to conveniently create, compare, and modulate comprehensive graph visualizations with a wide range of styles. We showcase various examples to verify the expressiveness of Nuwa. Via an expert interview and the analysis on cognitive dimensions we evaluate the usability of Nuwa.
Dongming Han, Wei Chen 0001, Jiacheng Pan, Xumeng Wang, Zhen Wen 0001, Luoxuan Weng, Minfeng Zhu 0001, Yingcai Wu, Rüdiger Westermann
PacificVis9
2024 GraphFederator: Federated Visual Analysis for Multi-party Graphs
abstract
This paper presents GraphFederator, a novel approach to construct federated representations of multi-party graphs and supports privacy-preserving visual analysis of graphs. Inspired by the concept of federated learning, we reformulate the analysis of multi-party graphs into a decentralization process. The new federation framework consists of a shared module that is responsible for federated modeling and analysis, and a set of local modules that run on respective graph data. Specifically, we propose a Federated Graph Representation Model (FGRM) that is learned from encrypted characteristics of multi-party graphs in local modules. We also design multiple visualization tools for federated visualization, exploration, and analysis of multi-party graphs. Experimental results on two datasets demonstrate the effectiveness of our approach.
Dongming Han, Wei Chen 0001, Rusheng Pan, Yijing Liu 0003, Jiehui Zhou, Haozhe Feng, Tian-Ye Zhang, Xumeng Wang, Minfeng Zhu 0001, Jianrong Tao, Changjie Fan, Xiaolong Zhang 0001
PacificVis10
2024 Exploring the neural landscape: Visual analytics of neuron activation in large language models with NeuronautLLM
abstract
Large language models (LLMs) like those that power OpenAI’s ChatGPT and Google’s Gemini have played a major part in the recent wave of machine learning and artificial intelligence advancements. However, interpreting LLMs and visualizing their components is extremely difficult due to the incredible scale and high dimensionality of model data. NeuronautLLM introduces a visual analysis system for identifying and visualizing influential neurons in transformer-based language models as they relate to user-defined prompts. Our approach combines simple, yet information-dense visualizations as well as neuron explanation and classification data to provide a wealth of opportunities for exploration. NeuronautLLM was reviewed by two experts to verify its efficacy as a tool for practical model interpretation. Interviews and usability tests with five LLM experts demonstrated NeuronautLLM’s exceptional usability and its readiness for real-world application. Furthermore, two in-depth case studies on model reasoning and social bias highlight NeuronautLLM’s versatility in aiding the analysis of a wide range of LLM research problems.
Ollie Woodman, Zhen Wen 0001, Yiwen Ren, Minfeng Zhu 0001, Wei Chen 0001
Graph. Model.5
2024 A visual analysis approach for data transformation via domain knowledge and intelligent models
Chengcan Chu, Minfeng Zhu 0001, Yating Wei, Jiacheng Pan, Dongming Han, Xuwei Tan, Wei Chen 0001
Multim. Syst.4
2024 PromptMagician: Interactive Prompt Engineering for Text-to-Image Creation
abstract
Generative text-to-image models have gained great popularity among the public for their powerful capability to generate high-quality images based on natural language prompts. However, developing effective prompts for desired images can be challenging due to the complexity and ambiguity of natural language. This research proposes PromptMagician, a visual analysis system that helps users explore the image results and refine the input prompts. The backbone of our system is a prompt recommendation model that takes user prompts as input, retrieves similar prompt-image pairs from DiffusionDB, and identifies special (important and relevant) prompt keywords. To facilitate interactive prompt refinement, PromptMagician introduces a multi-level visualization for the cross-modal embedding of the retrieved images and recommended keywords, and supports users in specifying multiple criteria for personalized exploration. Two usage scenarios, a user study, and expert interviews demonstrate the effectiveness and usability of our system, suggesting it facilitates prompt engineering and improves the creativity support of the generative text-to-image model.
Yingchaojie Feng, Xingbo Wang 0001, Kamkwai Wong, Yuhong Lu, Minfeng Zhu 0001, Baicheng Wang, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2024 Differentiable Design Galleries: A Differentiable Approach to Explore the Design Space of Transfer Functions
abstract
The transfer function is crucial for direct volume rendering (DVR) to create an informative visual representation of volumetric data. However, manually adjusting the transfer function to achieve the desired DVR result can be time-consuming and unintuitive. In this paper, we propose Differentiable Design Galleries, an image-based transfer function design approach to help users explore the design space of transfer functions by taking advantage of the recent advances in deep learning and differentiable rendering. Specifically, we leverage neural rendering to learn a latent design space, which is a continuous manifold representing various types of implicit transfer functions. We further provide a set of interactive tools to support intuitive query, navigation, and modification to obtain the target design, which is represented as a neural-rendered design exemplar. The explicit transfer function can be reconstructed from the target design with a differentiable direct volume renderer. Experimental results on real volumetric data demonstrate the effectiveness of our method.
Bo Pan 0004, Jiaying Lu 0005, Weifeng Chen 0003, Yiyao Wang, Minfeng Zhu 0001, Chenhao Yu, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2024 Quantivine: A Visualization Approach for Large-Scale Quantum Circuit Representation and Analysis
abstract
Quantum computing is a rapidly evolving field that enables exponential speed-up over classical algorithms. At the heart of this revolutionary technology are quantum circuits, which serve as vital tools for implementing, analyzing, and optimizing quantum algorithms. Recent advancements in quantum computing and the increasing capability of quantum devices have led to the development of more complex quantum circuits. However, traditional quantum circuit diagrams suffer from scalability and readability issues, which limit the efficiency of analysis and optimization processes. In this research, we propose a novel visualization approach for large-scale quantum circuits by adopting semantic analysis to facilitate the comprehension of quantum circuits. We first exploit meta-data and semantic information extracted from the underlying code of quantum circuits to create component segmentations and pattern abstractions, allowing for easier wrangling of massive circuit diagrams. We then develop Quantivine, an interactive system for exploring and understanding quantum circuits. A series of novel circuit visualizations is designed to uncover contextual details such as qubit provenance, parallelism, and entanglement. The effectiveness of Quantivine is demonstrated through two usage scenarios of quantum circuits with up to 100 qubits and a formal user evaluation with quantum experts. A free copy of this paper and all supplemental materials are available at https://osf.io/2m9yh/?view_only=0aa1618c97244f5093cd7ce15f1431f9.
Zhen Wen 0001, Siwei Tan, Jieyi Chen, Minfeng Zhu 0001, Dongming Han, Jianwei Yin, Mingliang Xu 0001, Wei Chen 0001
IEEE Trans. Vis. Comput. Graph.5
2024 A Parallel Framework for Streaming Dimensionality Reduction
abstract
The visualization of streaming high-dimensional data often needs to consider the speed in dimensionality reduction algorithms, the quality of visualized data patterns, and the stability of view graphs that usually change over time with new data. Existing methods of streaming high-dimensional data visualization primarily line up essential modules in a serial manner and often face challenges in satisfying all these design considerations. In this research, we propose a novel parallel framework for streaming high-dimensional data visualization to achieve high data processing speed, high quality in data patterns, and good stability in visual presentations. This framework arranges all essential modules in parallel to mitigate the delays caused by module waiting in serial setups. In addition, to facilitate the parallel pipeline, we redesign these modules with a parametric non-linear embedding method for new data embedding, an incremental learning method for online embedding function updating, and a hybrid strategy for optimized embedding updating. We also improve the coordination mechanism among these modules. Our experiments show that our method has advantages in embedding speed, quality, and stability over other existing methods to visualize streaming high-dimensional data.
Jiazhi Xia, Linquan Huang, Yiping Sun, Zhiwei Deng, Xiaolong Zhang 0001, Minfeng Zhu 0001
IEEE Trans. Vis. Comput. Graph.6
2023 A privacy-aware visual query approach for location-based data
Ziliang Wu, Erqing Zhang, Zhaosong Huang, Mingliang Xu 0001, Lechao Cheng, Minfeng Zhu 0001, Wei Chen 0001
Comput. Graph.7
2023 Interactive visual analytics of parallel training strategies for DNN models
Yating Wei, Gongchang Ou, Han Gao 0016, Minfeng Zhu 0001, Wei Chen 0001
Comput. Graph.8
2022 Interactive Image Synthesis with Panoptic Layout Generation
abstract
Interactive image synthesis from user-guided input is a challenging task when users wish to control the scene structure of a generated image with ease. Although remarkable progress has been made on layout-based image synthesis approaches, existing methods require high-precision inputs such as accurately placed bounding boxes, which might be constantly violated in an interactive setting. When placement of bounding boxes is subject to perturbation, layout-based models suffer from “missing regions” in the constructed semantic layouts and hence undesirable artifacts in the generated images. In this work, we propose Panoptic Layout Generative Adversarial Network (PLGAN) to address this challenge. The PLGAN employs panoptic theory which distinguishes object categories between “stuff” with amorphous boundaries and “things” with well-defined shapes, such that stuff and instance layouts are constructed through separate branches and later fused into panoptic layouts. In particular, the stuff layouts can take amorphous shapes and fill up the missing regions left out by the instance layouts. We experimentally compare our PLGAN with state-of-the-art layout-based models on the COCO-Stuff, Visual Genome, and Landscape datasets. The advantages of PLGAN are not only visually demonstrated but quantitatively verified in terms of inception score, Fréchet inception distance, classification accuracy score, and coverage. The code is available at https://github.com/wb-finalking/PLGAN.
Bo Wang 0071, Minfeng Zhu 0001
CVPR3
2021 SHOT-VAE: Semi-supervised Deep Generative Models With Label-aware ELBO Approximations
abstract
Semi-supervised variational autoencoders (VAEs) have obtained strong results, but have also encountered the challenge that good ELBO values do not always imply accurate inference results.In this paper, we investigate and propose two causes of this problem: (1) The ELBO objective cannot utilize the label information directly. (2) A bottleneck value exists, and continuing to optimize ELBO after this value will not improve inference accuracy. On the basis of the experiment results, we propose SHOT-VAE to address these problems without introducing additional prior knowledge. The SHOT-VAE offers two contributions: (1) A new ELBO approximation named smooth-ELBO that integrates the label predictive loss into ELBO. (2) An approximation based on optimal interpolation that breaks the ELBO value bottleneck by reducing the margin between ELBO and the data likelihood. The SHOT-VAE achieves good performance with 25.30% error rate on CIFAR-100 with 10k labels and reduces the error rate to 6.11% on CIFAR-10 with 4k labels.
Haozhe Feng, Kezhi Kong, Tian-Ye Zhang, Minfeng Zhu 0001, Wei Chen 0001
AAAI5
2021 KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation
abstract
Conventional unsupervised multi-source domain adaptation (UMDA) methods assume all source domains can be accessed directly. However, this assumption neglects the privacy-preserving policy, where all the data and computations must be kept decentralized. There exist three challenges in this scenario: (1) Minimizing the domain distance requires the pairwise calculation of the data from the source and target domains, while the data on the source domain is not available. (2) The communication cost and privacy security limit the application of existing UMDA methods, such as the domain adversarial training. (3) Since users cannot govern the data quality, the irrelevant or malicious source domains are more likely to appear, which causes negative transfer. To address the above problems, we propose a privacy-preserving UMDA paradigm named Knowledge Distillation based Decentralized Domain Adaptation (KD3A), which performs domain adaptation through the knowledge distillation on models from different source domains. The extensive experiments show that KD3A significantly outperforms state-of-the-art UMDA approaches. Moreover, the KD3A is robust to the negative transfer and brings a 100x reduction of communication cost compared with other decentralized UMDA methods.
Haozhe Feng, Zhaoyang You, Tian-Ye Zhang, Minfeng Zhu 0001, Fei Wu 0001, Chao Wu 0001, Wei Chen 0001
ICML5
2021 Exemplar-based Layout Fine-tuning for Node-link Diagrams
abstract
We design and evaluate a novel layout fine-tuning technique for node-link diagrams that facilitates exemplar-based adjustment of a group of substructures in batching mode. The key idea is to transfer user modifications on a local substructure to other substructures in the entire graph that are topologically similar to the exemplar. We first precompute a canonical representation for each substructure with node embedding techniques and then use it for on-the-fly substructure retrieval. We design and develop a light-weight interactive system to enable intuitive adjustment, modification transfer, and visual graph exploration. We also report some results of quantitative comparisons, three case studies, and a within-participant user study.
Jiacheng Pan, Wei Chen 0001, Shuyue Zhou, Wei Zeng 0004, Minfeng Zhu 0001, Jian Chen 0006, Siwei Fu, Yingcai Wu
IEEE Trans. Vis. Comput. Graph.6
2021 DRGraph: An Efficient Graph Layout Algorithm for Large-scale Graphs by Dimensionality Reduction
abstract
Efficient layout of large-scale graphs remains a challenging problem: the force-directed and dimensionality reduction-based methods suffer from high overhead for graph distance and gradient computation. In this paper, we present a new graph layout algorithm, called DRGraph, that enhances the nonlinear dimensionality reduction process with three schemes: approximating graph distances by means of a sparse distance matrix, estimating the gradient by using the negative sampling technique, and accelerating the optimization process through a multi-level layout scheme. DRGraph achieves a linear complexity for the computation and memory consumption, and scales up to large-scale graphs with millions of nodes. Experimental results and comparisons with state-of-the-art graph layout methods demonstrate that DRGraph can generate visually comparable layouts with a faster running time and a lower memory requirement.
Minfeng Zhu 0001, Wei Chen 0001, Yuxuan Hou, Liangjun Liu, Kaiyuan Zhang 0002
IEEE Trans. Vis. Comput. Graph.1
2020 EEMEFN: Low-Light Image Enhancement via Edge-Enhanced Multi-Exposure Fusion Network
abstract
This work focuses on the extremely low-light image enhancement, which aims to improve image brightness and reveal hidden information in darken areas. Recently, image enhancement approaches have yielded impressive progress. However, existing methods still suffer from three main problems: (1) low-light images usually are high-contrast. Existing methods may fail to recover images details in extremely dark or bright areas; (2) current methods cannot precisely correct the color of low-light images; (3) when the object edges are unclear, the pixel-wise loss may treat pixels of different objects equally and produce blurry images. In this paper, we propose a two-stage method called Edge-Enhanced Multi-Exposure Fusion Network (EEMEFN) to enhance extremely low-light images. In the first stage, we employ a multi-exposure fusion module to address the high contrast and color bias issues. We synthesize a set of images with different exposure time from a single image and construct an accurate normal-light image by combining well-exposed areas under different illumination conditions. Thus, it can produce realistic initial images with correct color from extremely noisy and low-light images. Secondly, we introduce an edge enhancement module to refine the initial images with the help of the edge information. Therefore, our method can reconstruct high-quality images with sharp edges when minimizing the pixel-wise loss. Experiments on the See-in-the-Dark dataset indicate that our EEMEFN approach achieves state-of-the-art performance.
Minfeng Zhu 0001, Pingbo Pan, Wei Chen 0001, Yi Yang 0001
AAAI1
2020 A Natural-language-based Visual Query Approach of Uncertain Human Trajectories
abstract
Visual querying is essential for interactively exploring massive trajectory data. However, the data uncertainty imposes profound challenges to fulfill advanced analytics requirements. On the one hand, many underlying data does not contain accurate geographic coordinates, e.g., positions of a mobile phone only refer to the regions (i.e., mobile cell stations) in which it resides, instead of accurate GPS coordinates. On the other hand, domain experts and general users prefer a natural way, such as using a natural language sentence, to access and analyze massive movement data. In this paper, we propose a visual analytics approach that can extract spatial-temporal constraints from a textual sentence and support an effective query method over uncertain mobile trajectory data. It is built up on encoding massive, spatially uncertain trajectories by the semantic information of the POls and regions covered by them, and then storing the trajectory documents in text database with an effective indexing scheme. The visual interface facilitates query condition specification, situation-aware visualization, and semantic exploration of large trajectory data. Usage scenarios on real-world human mobility datasets demonstrate the effectiveness of our approach.
Zhaosong Huang, Ye Zhao 0003, Wei Chen 0001, Shengjie Gao, Kejie Yu, MingJie Tang, Minfeng Zhu 0001, Mingliang Xu 0001
IEEE Trans. Vis. Comput. Graph.8
2019 DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-To-Image Synthesis
abstract
In this paper, we focus on generating realistic images from text descriptions. Current methods first generate an initial image with rough shape and color, and then refine the initial image to a high-resolution one. Most existing text-to-image synthesis methods have two main problems. (1) These methods depend heavily on the quality of the initial images. If the initial image is not well initialized, the following processes can hardly refine the image to a satisfactory quality. (2) Each word contributes a different level of importance when depicting different image contents, however, unchanged text representation is used in existing image refinement processes. In this paper, we propose the Dynamic Memory Generative Adversarial Network (DM-GAN) to generate high-quality images. The proposed method introduces a dynamic memory module to refine fuzzy image contents, when the initial images are not well generated. A memory writing gate is designed to select the important text information based on the initial image content, which enables our method to accurately generate images from the text description. We also utilize a response gate to adaptively fuse the information read from the memories and the image features. We evaluate the DM-GAN model on the Caltech-UCSD Birds 200 dataset and the Microsoft Common Objects in Context dataset. Experimental results demonstrate that our DM-GAN model performs favorably against the state-of-the-art approaches.
Minfeng Zhu 0001, Pingbo Pan, Wei Chen 0001, Yi Yang 0001
CVPR1
2019 FBVA: A Flow-Based Visual Analytics Approach for Citywide Crowd Mobility
abstract
Analyzing structures of crowd mobility at city level is a challenging task due to the complex crowd mobility and dynamic changes generated by the social activities over time. These structures, defined as high-dimensional mobility structures (HMSs), contain spatiotemporal information and are simultaneously influenced by the geographical distributions and daily activities of citywide crowd. However, few work has been dedicated to depict and analyze these structures, mainly due to the lack of effective models. In this paper, we propose to model the crowd mobility as a dynamical system and characterize the irregular mobility data with a novel local coherence of sparse field (LCSF) algorithm. The proposed algorithm makes it possible to measure the separation behavior of trajectories in an irregular and sparse topology network. Detected HMS, referred as local separation measure of LCSF, divides the geographical urban areas into distinct functional regions over time. We design and implement a visual analytics system to facilitate situation-aware analysis of a huge amount of crowd mobility and their socialized behaviors. Case studies based on a real-world data set demonstrate the effectiveness of the proposed approach.
Yuan Yuan 0012, Minfeng Zhu 0001, Liang Chang 0003, Xiyan Sun, Zi'ang Ding
IEEE Trans. Comput. Soc. Syst.4
2019 Location2vec: A Situation-Aware Representation for Visual Exploration of Urban Locations
abstract
Understanding the relationship between urban locations is an essential task in urban planning and transportation management. Although prior works have focused on studying urban locations by aggregating location-based properties, our scheme preserves the mutual influence between urban locations and mobility behavior, and thereby enables situation-aware exploration of urban regions. By leveraging word embedding techniques, we encode urban locations with a vectorized representation while retaining situational awareness. Specifically, we design a spatial embedding algorithm that is precomputed by incorporating the interactions between urban locations and moving objects. To explore our proposed technique, we have designed and implemented a web-based visual exploration system that supports the comprehensive analysis of human mobility, location functionality, and traffic assessment by leveraging the proposed visual representation. The case studies demonstrate the effectiveness of our approach.
Minfeng Zhu 0001, Wei Chen 0001, Jiazhi Xia, Yuxin Ma 0001, Yankong Zhang, Yuetong Luo, Zhaosong Huang, Liangjun Liu
IEEE Trans. Intell. Transp. Syst.1
2018 Structuring Mobility Transition With an Adaptive Graph Representation
abstract
Modeling human mobility is a critical task in fields such as urban planning, ecology, and epidemiology. Given the current use of mobile phones, there is an abundance of data that can be used to create models of high reliability. Existing techniques can reveal the macropatterns of crowd movement or analyze the trajectory of a person; however, they typically focus on geographical characteristics. This paper presents a graph-based approach for structuring crowd mobility transition over multiple granularities in the context of social behavior. The key to our approach is an adaptive data representation, the adaptive mobility transition graph (AMTG), which is globally generated from citywide human mobility data by defining the temporal trends of human mobility and the interleaved transitions between different mobility patterns. We describe the design, creation, and manipulation of the AMTG and introduce a visual analysis system that supports the multifaceted exploration of citywide human mobility patterns.
Tianlong Gu, Minfeng Zhu 0001, Wei Chen 0001, Zhaosong Huang, Ross Maciejewski, Liang Chang 0003
IEEE Trans. Comput. Soc. Syst.2
2018 VAUD: A Visual Analysis Approach for Exploring Spatio-Temporal Urban Data
abstract
Urban data is massive, heterogeneous, and spatio-temporal, posing a substantial challenge for visualization and analysis. In this paper, we design and implement a novel visual analytics approach, Visual Analyzer for Urban Data (VAUD), that supports the visualization, querying, and exploration of urban data. Our approach allows for cross-domain correlation from multiple data sources by leveraging spatial-temporal and social inter-connectedness features. Through our approach, the analyst is able to select, filter, aggregate across multiple data sources and extract information that would be hidden to a single data subset. To illustrate the effectiveness of our approach, we provide case studies on a real urban dataset that contains the cyber-, physical-, and social- information of 14 million citizens over 22 days.
Wei Chen 0001, Zhaosong Huang, Feiran Wu, Minfeng Zhu 0001, Huihua Guan, Ross Maciejewski
IEEE Trans. Vis. Comput. Graph.4