VLDB 2026 Research / reviewers in the wild / expert
Linzhuang Sun
dblp:353/0971
· DBLP profile ↗
16ranked-venue papers
1as first author
16since 2021 · last 2025
0000-0003-0363-3607ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought VerificationabstractLinzhuang Sun, Hao Liang, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang, Zenan Zhou, Wentao Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Linzhuang Sun, Hao Liang 0017, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang 0132, Zenan Zhou, Wentao Zhang 0001 |
ACL (1) | 1 |
| 2025 | From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and EditingabstractWe introduce the task of text-to-diagram generation, which focuses on creating structured visual representations directly from textual descriptions. Existing approaches in text-to-image and text-to-code generation lack the logical organization and flexibility needed to produce accurate, editable diagrams, often resulting in outputs that are either unstructured or difficult to modify. To address this gap, we introduce DiagramGenBenchmark, a comprehensive evaluation framework encompassing eight distinct diagram categories, including flowcharts, model architecture diagrams, and mind maps. Additionally, we present DiagramAgent, an innovative framework with four core modules—Plan Agent, Code Agent, Check Agent, and Diagram-to-Code Agent—designed to facilitate both the generation and refinement of complex diagrams. Our extensive experiments, which combine objective metrics with human evaluations, demonstrate that DiagramAgent significantly outperforms existing baseline models in terms of accuracy, structural coherence, and modifiability. This work not only establishes a foundational benchmark for the text-to-diagram generation task but also introduces a powerful toolset to advance research and applications in this emerging area. Jingxuan Wei, Cheng Tan 0012, Siyuan Li 0002, Zhangyang Gao, Linzhuang Sun, Bihui Yu, Ruifeng Guo |
CVPR | 7 |
| 2025 | ChiImpAVE: An Open-Source Benchmark for Chinese Implicit Attribute Value Extraction
Bihui Yu, Huiyang Shi, Linzhuang Sun, Jingxuan Wei |
ICIC (18) | 7 |
| 2025 | Beyond Relevance: Utility-Driven Retrieval for Visual Document Question Answering
Bihui Yu, Zhuoya Yao, Huiyang Shi, Liping Bu, Linzhuang Sun, Jingxuan Wei |
ICIC (16) | 7 |
| 2025 | EEGTCT: Electroencephalogram-Based Chinese Text Decoding
Bihui Yu, Jingxuan Wei, Linzhuang Sun, Liping Bu |
ICIC (21) | 6 |
| 2025 | MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical ContextsabstractWith the rapid progress of Multimodal LLMs, evaluating their mathematical reasoning capabilities has become an increasingly important research direction. In particular, visual-textual mathematical reasoning serves as a key indicator of an MLLM's ability to comprehend and solve complex, multi-step quantitative problems. While existing benchmarks such as MathVista and MathVerse have advanced the evaluation of multimodal math proficiency, they primarily rely on digitally rendered content and fall short in capturing the complexity of real-world scenarios. To bridge this gap, we introduce MathScape, a novel benchmark focused on assessing MLLMs' reasoning ability in realistic mathematical contexts. MathScape comprises 1,369 high-quality math problems paired with human-captured real-world images, closely reflecting the challenges encountered in practical educational settings. We conduct a thorough multi-dimensional evaluation across nine leading closed-source MLLMs, three open-source MLLMs with over 20 billion parameters, and seven smaller-scale MLLMs. Our results show that even SOTA models struggle with real-world math tasks, lagging behind human performance-highlighting critical limitations in current model capabilities. Moreover, we find that strong performance on synthetic or digitally rendered images does not guarantee similar effectiveness on real-world tasks. This underscores the necessity of MathScape in the next stage of multimodal mathematical reasoning. Hao Liang 0017, Linzhuang Sun, zhouminxuan zhouminxuan, Zirong Chen, Meiyi Qiang, Tianpeng Li, Fan Yang 0132, Zenan Zhou, Wentao Zhang 0001 |
ACM Multimedia | 2 |
| 2025 | ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningabstractLarge Language Models (LLMs) have shown remarkable capabilities in reasoning, exemplified by the success of OpenAI-o1 and DeepSeek-R1. However, integrating reasoning with external search processes remains challenging, especially for complex multi-hop questions requiring multiple retrieval steps. We propose ReSearch, a novel framework that trains LLMs to Reason with Search via reinforcement learning without using any supervised data on reasoning steps. Our approach treats search operations as integral components of the reasoning chain, where when and how to perform searches is guided by text-based thinking, and search results subsequently influence further reasoning. We train ReSearch on Qwen2.5-7B(-Instruct) and Qwen2.5-32B(-Instruct) models and conduct extensive experiments. Despite being trained on only one dataset, our models demonstrate strong generalizability across various benchmarks. Analysis reveals that ReSearch naturally elicits advanced reasoning capabilities such as reflection and self-correction during the reinforcement learning process. Mingyang Chen 0002, Linzhuang Sun, Tianpeng Li, Haoze Sun, Chenzheng Zhu, Haofen Wang, Jeff Z. Pan, Wen Zhang 0015, Huajun Chen, Fan Yang 0132, Zenan Zhou, Weipeng Chen |
NeurIPS | 2 |
| 2025 | BRACE: A Benchmark for Robust Audio Caption Quality EvaluationabstractAutomatic audio captioning is essential for audio understanding, enabling applications such as accessibility and content indexing. However, evaluating the quality of audio captions remains a major challenge, especially in reference-free settings where high-quality ground-truth captions are unavailable. While CLAPScore is currently the most widely used reference-free Audio Caption Evaluation Metric(ACEM), its robustness under diverse conditions has not been systematically validated. To address this gap, we introduce BRACE, a new benchmark designed to evaluate audio caption alignment quality in a reference-free setting. BRACE is primarily designed for assessing ACEMs, and can also be extended to measure the modality alignment abilities of Large Audio Language Model(LALM). BRACE consists of two sub-benchmarks: BRACE-Main for fine-grained caption comparison and BRACE-Hallucination for detecting subtle hallucinated content. We construct these datasets through high-quality filtering, LLM-based corruption, and human annotation. Given the widespread adoption of CLAPScore as a reference-free ACEM and the increasing application of LALMs in audio-language tasks, we evaluate both approaches using the BRACE benchmark, testing CLAPScore across various CLAP model variants and assessing multiple LALMs. Notably, even the best-performing CLAP-based ACEM achieves only a 70.01 F1-score on the BRACE-Main benchmark, while the best LALM reaches just 63.19. By revealing the limitations of CLAP models and LALMs, our BRACE benchmark offers valuable insights into the direction of future research. Our evaluation code and benchmark dataset are released in https://github.com/HychTus/BRACEEvaluation and https://huggingface.co/datasets/gtysssp/audiobenchmarks. Tianyu Guo 0001, Hao Liang 0017, Meiyi Qiang, Bohan Zeng, Linzhuang Sun, Bin Cui 0001, Wentao Zhang 0001 |
NeurIPS | 6 |
| 2025 | Brain-inspired computing based on deep learning for human-computer interaction: A review
Bihui Yu, Jingxuan Wei, Linzhuang Sun, Liping Bu |
Neurocomputing | 5 |
| 2024 | Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-consistency Training
Cheng Tan 0012, Jingxuan Wei, Zhangyang Gao, Linzhuang Sun, Siyuan Li 0002, Ruifeng Guo, Bihui Yu, Stan Z. Li |
ECCV (40) | 4 |
| 2024 | Sentence-Level or Token-Level? A Comprehensive Study on Knowledge Distillation
Jingxuan Wei, Linzhuang Sun, Yichong Leng, Xu Tan 0003, Bihui Yu, Ruifeng Guo |
IJCAI | 2 |
| 2024 | Interpretable and Generalizable Spatiotemporal Predictive Learning with Disentangled Consistency
Jingxuan Wei, Cheng Tan 0012, Zhangyang Gao, Linzhuang Sun, Bihui Yu, Ruifeng Guo, Stan Z. Li |
ECML/PKDD (3) | 4 |
| 2024 | SAM-Wav2lip++: Enhancing Behavioral Realism in Synthetic Agents Through Audio-Driven Speech and Action RefinementabstractDigital human generation is a forward-looking field in technology. Despite significant progress in the generation of speaking facial videos, many challenges remain unaddressed. Issues such as unnatural head movements, distorted expressions, artifacts in generated videos, and uncoordinated limb movements persist. Most current efforts are focused on specific individuals, with enhancements often limited to head movements without further advancing the overall behavioral actions of digital humans. In this context, we introduce a new dataset, CFMD, and a novel model, SAM-Wav2lip++, capable of generating consistent, audio-synchronized lip and behavior action videos from a single reference image of any identity. This work features three main innovative components: (1) a contrastive lip-sync discriminator for precise lip synchronization, (2) a generator for the synthesis of sound-action consistency, and (3) the SAM module for facial refinement operations. Through extensive experiments and user studies, our results demonstrate that our model can synthesize digital human videos of impressively high perceptual quality that accurately sync lip movements and behavioral actions with the input audio, substantially outperforming the state-of-the-art baselines evaluations. Bihui Yu, Huiyang Shi, Guiyong Chang, Jingxuan Wei, Linzhuang Sun, Songtao Tian, Liping Bu |
SMC | 6 |
| 2024 | Faster and More Efficient Subject Image Generation for Text-to-Image Diffusion ModelsabstractIn recent years, there has been significant progress in text-to-image generation models. However, text struggles to accurately describe abstract concepts like shapes and sizes. Some methods have been proposed to enhance text prompt by incorporating image prompts. While they have shown effective improvements, they either require substantial fine-tuning costs or struggle to effectively integrate text and image information. In our study, we delve into the issue of the difficulty in integrating text and image information in decoupled cross-attention and conduct visual analysis. We identify the presence of background-related tokens in image features as a key factor affecting text fidelity. To address this issue, we develop an algorithm to filter out these tokens. Additionally, we observe differences in the attention of Unet layers to text prompts and image prompts. Based on this finding, we optimize the flow of image information to reduce interference with text information. In summary, we introduce a new topic-customized method that requires no repeated training. It trains a plug-and-play image prompt adapter with only 417M parameters, lightweight yet powerful, surpassing existing models in both text and image consistency. Our code and pre-trained checkpoints will be available at https://github.com/YZBPXX/DDCA. Bihui Yu, Zhengbing Yao, Jingxuan Wei, Linzhuang Sun, Liping Bu |
SMC | 4 |
| 2024 | Enhancing human-like multimodal reasoning: a new challenging dataset and comprehensive framework
Jingxuan Wei, Cheng Tan 0012, Zhangyang Gao, Linzhuang Sun, Siyuan Li 0002, Bihui Yu, Ruifeng Guo, Stan Z. Li |
Neural Comput. Appl. | 4 |
| 2023 | TED-CS: Textual Enhanced Sensitive Video Detection with Common Sense Knowledge
Bihui Yu, Linzhuang Sun, Jingxuan Wei, Shuyue Tan, Yiman Zhao, Liping Bu |
ADMA (2) | 2 |