VLDB 2026 Research / reviewers in the wild / expert
Yiyao Wang
dblp:273/0668
· DBLP profile ↗
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic FrameworkabstractVisualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep research and generate comprehensive reports. Despite its progress, existing deep research frameworks primarily focus on generating text-only content, leaving the automated generation of interleaved texts and visualizations underexplored. This novel task poses key challenges in designing informative visualizations and effectively integrating them with text reports. To address these challenges, we propose Formal Description of Visualization (FDV), a structured textual representation of charts that enables LLMs to learn from and generate diverse, high-quality visualizations. Building on this representation, we introduce Multimodal DeepResearcher, an agentic framework that decomposes the task into four stages: (1) researching, (2) exemplar report textualization, (3) planning and (4) multimodal report generation. For the evaluation of the generated reports, we develop MultimodalReportBench which contains 100 diverse topics as inputs, and a set of dedicated metrics for report and chart evaluation. Extensive experiments across models and evaluation methods demonstrate the effectiveness of Multimodal DeepResearcher. Notably, utilizing the same Claude 3.7 Sonnet model, Multimodal DeepResearcher achieves an 82% overall win rate over the baseline method. Zhaorui Yang 0001, Bo Pan 0004, Yiyao Wang, Xingyu Liu 0003, Luoxuan Weng, Yingchaojie Feng, Haozhe Feng, Minfeng Zhu 0001, Wei Chen 0001 |
AAAI | 4 |
| 2026 | IGenBench: Benchmarking the Reliability of Text-to-Infographic GenerationabstractYinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan, Yupeng Xie, Jiale Lao, Yiyao Wang, Haoxuan Li, Tingting Gao, Bo Pan, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu, Yingchaojie Feng, Yuyu Luo, Wei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yinghao Tang, Xueding Liu, Tingfeng Lan, Jiale Lao, Yiyao Wang, Tingting Gao, Bo Pan 0004, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu 0001, Yingchaojie Feng, Yuyu Luo, Wei Chen 0001 |
ACL (1) | 7 |
| 2026 | ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language ModelsabstractLarge language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks. Understanding how LLMs internally represent knowledge remains a significant challenge. Despite Sparse Autoencoders (SAEs) have emerged as a promising technique for extracting interpretable features from LLMs, SAE features do not inherently align with human-understandable concepts, making their interpretation cumbersome and labor-intensive. To bridge the gap between SAE features and human concepts, we present ConceptViz, a visual analytics system designed for exploring concepts in LLMs. ConceptViz implements a novel Identification ⇒ Interpretation ⇒Validation pipeline, enabling users to query SAEs using concepts of interest, interactively explore concept-to-feature alignments, and validate the correspondences through model behavior verification. We demonstrate the effectiveness of ConceptViz through two usage scenarios and a user study. Our results show that ConceptViz enhances interpretability research by streamlining the discovery and validation of meaningful concept representations in LLMs, ultimately aiding researchers in building more accurate mental models of LLM features. Our code and user guide are publicly available at https://github.com/Happy-Hippo209/conceptViz. Zhen Wen 0001, Qiqi Jiang, Chenxiao Li, Yiyao Wang, Xiuqi Huang, Minfeng Zhu 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2026 | VidGuard3D: A Visual Risk Analysis Approach for Protecting 3D Assets Against Video-Based Reconstruction AttacksabstractThe unauthorized acquisition of 3D assets by means of advanced techniques in 3D reconstruction is a major but often overlooked threat to publishers of videos. Preventing such threats is challenging due to the uninterpretable nature of 3D reconstruction and the diversity in requirements of 3D model demonstration. In this paper, we introduce VidGuard3D-a visual risk analysis approach that quantifies and locates the sources of risk for video-based 3D asset reconstruction attacks. Our approach uses attack simulation to support users in formulating a comprehensive understanding of 3D asset leakage risks, with a particular focus on the correlations between video segments and exposure risks of user-specified areas in the asset. We also proposed a prototype system that integrates this approach to facilitate video editing according to the knowledge of correlations. Two operations of video editing can be swiftly applied by users to form editing plans and minimize detected leakage risks. Finally, we conducted a user study and case studies that demonstrated the practicality and effectiveness of our approach. Yiyao Wang, Ollie Woodman, Shenghui Hu, Ruizhe Pan, Bo Pan 0004, Xumeng Wang, Minfeng Zhu 0001, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Function-Centric Bayesian Network for Zero-Shot Object Goal Navigation
Sixian Zhang, Xinyao Yu 0002, Xinhang Song, Yiyao Wang, Shuqiang Jiang |
ICCV | 4 |
| 2025 | An opinion leader mining method based on text contents and network features
Yiyao Wang, Zurui Gan, Tiejun Xi, Jianqing Xi, Xiran Xu, Deyu Qi 0001 |
World Wide Web (WWW) | 2 |
| 2024 | Differentiable Design Galleries: A Differentiable Approach to Explore the Design Space of Transfer FunctionsabstractThe transfer function is crucial for direct volume rendering (DVR) to create an informative visual representation of volumetric data. However, manually adjusting the transfer function to achieve the desired DVR result can be time-consuming and unintuitive. In this paper, we propose Differentiable Design Galleries, an image-based transfer function design approach to help users explore the design space of transfer functions by taking advantage of the recent advances in deep learning and differentiable rendering. Specifically, we leverage neural rendering to learn a latent design space, which is a continuous manifold representing various types of implicit transfer functions. We further provide a set of interactive tools to support intuitive query, navigation, and modification to obtain the target design, which is represented as a neural-rendered design exemplar. The explicit transfer function can be reconstructed from the target design with a differentiable direct volume renderer. Experimental results on real volumetric data demonstrate the effectiveness of our method. Bo Pan 0004, Jiaying Lu 0005, Weifeng Chen 0003, Yiyao Wang, Minfeng Zhu 0001, Chenhao Yu, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Joint Resource Allocation and Passive Beamforming in RIS-Aided HetNets With Wireless BackhaulabstractHeterogeneous networks(HetNets) have been widely used in the development of 5G because a large number ofsmall base stations(SBSs) can be deployed in hot spots to alleviate uneven traffic distribution. However, the performance of HetNet with wireless backhaul is limited by its backhaul capacity and the severe wireless interference. To solve this problem, we introducereconfigurable intelligent surface(RIS) into the wireless backhaul HetNets and employ reversedtime division duplex(TDD) in the time domain and dynamicsoft frequency reuse(SFR) in the frequency domain. In the proposed RIS-aided dual-layer HetNets system, we formulate a joint optimization problem of bandwidth allocation, RIS passive beamforming, and power allocation to maximize the overall data rate. Two schemes are considered to adapt to different situations, i.e.,unified dynamic SFR(U-SFR) andindividual dynamic SFR(I-SFR). We solve the different sub-problems in U-SFR and I-SFR modes with the alternate optimization method, and obtain closed-form expressions for bandwidth allocation and power allocation. The convergence, feasibility, and complexity of the proposed schemes are also analyzed. In numerical results, the overall data rate is shown to be significantly increased with the help of RIS. Jinping Niu, Yiyao Wang, Xiangwei Zhou |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | Classifying Facial Regions for Face HallucinationabstractRecently, convolutional neural networks (CNNs) have dominated the face hallucination task due to their powerful feature representation capability. However, most of them simply use the same weights to treat different facial regions without considering the reconstruction difficulty of different facial regions, resulting in the component regions (e.g., eyes, nose, mouth) of the reconstructed faces tending to be blurred. In this paper, we propose a novel facial region classification network (FRCN) to address this problem. The proposed method first divides the input low-resolution (LR) facial image into several patch blocks, then classifies them into three categories according to their reconstruction difficulty, and finally inputs the three types of patch blocks into three networks with different weights for reconstruction and combining, thereby recovering high-quality high-resolution (HR) facial image. Experimental results show that FRCN can remarkably improve face reconstruction's performance. Yiyao Wang, Tao Lu 0001, Yuanzhi Wang, Zhongyuan Wang 0001 |
IEEE Signal Process. Lett. | 1 |