Bihui Yu

dblp:323/7684 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 GenProve: Learning to Generate Text with Fine-Grained Provenance
abstract
Jingxuan Wei, Xingyue Wang, Yanghaoyu Liao, Jie Dong, Yuchen Liu, Caijun Jia, Bihui Yu, Junnan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingxuan Wei, Yanghaoyu Liao, Caijun Jia, Bihui Yu, Junnan Zhu
ACL (1)7
2026 Task-Aligned Crystallographic Descriptor Supervision for Polymorph Stability Ranking with Graph Neural Networks
Bihui Yu, Honghao He, Jingxuan Wei
ICIC (9)1
2026 mChartQA and mChartQABench: A multimodal-only solution for complex chart question-answering
Jingxuan Wei, Nan Xu 0004, Guiyong Chang, Yin Luo, Bihui Yu, Ruifeng Guo
Pattern Recognit.5
2025 MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
abstract
Linzhuang Sun, Hao Liang, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang, Zenan Zhou, Wentao Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Linzhuang Sun, Hao Liang 0017, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang 0132, Zenan Zhou, Wentao Zhang 0001
ACL (1)4
2025 From Words to Structured Visuals: A Benchmark and Framework for Text-to-Diagram Generation and Editing
abstract
We introduce the task of text-to-diagram generation, which focuses on creating structured visual representations directly from textual descriptions. Existing approaches in text-to-image and text-to-code generation lack the logical organization and flexibility needed to produce accurate, editable diagrams, often resulting in outputs that are either unstructured or difficult to modify. To address this gap, we introduce DiagramGenBenchmark, a comprehensive evaluation framework encompassing eight distinct diagram categories, including flowcharts, model architecture diagrams, and mind maps. Additionally, we present DiagramAgent, an innovative framework with four core modules—Plan Agent, Code Agent, Check Agent, and Diagram-to-Code Agent—designed to facilitate both the generation and refinement of complex diagrams. Our extensive experiments, which combine objective metrics with human evaluations, demonstrate that DiagramAgent significantly outperforms existing baseline models in terms of accuracy, structural coherence, and modifiability. This work not only establishes a foundational benchmark for the text-to-diagram generation task but also introduces a powerful toolset to advance research and applications in this emerging area.
Jingxuan Wei, Cheng Tan 0012, Siyuan Li 0002, Zhangyang Gao, Linzhuang Sun, Bihui Yu, Ruifeng Guo
CVPR8
2025 ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering
abstract
Chart question answering (CQA) has become a critical multimodal task for evaluating the reasoning capabilities of vision-language models.While early approaches have shown promising performance by focusing on visual features or leveraging large-scale pre-training, most existing evaluations rely on rigid output formats and objective metrics, thus ignoring the complex, real-world demands of practical chart analysis.In this paper, we introduce ChartMind, a new benchmark designed for complex CQA tasks in real-world settings.ChartMind covers seven task categories, incorporates multilingual contexts, supports open-domain textual outputs, and accommodates diverse chart formats, bridging the gap between real-world applications and traditional academic benchmarks.Furthermore, we propose a context-aware yet modelagnostic framework, ChartLLM, that focuses on extracting key contextual elements, reducing noise, and enhancing the reasoning accuracy of multimodal large language models.Extensive evaluations on ChartMind and three representative public benchmarks with 14 mainstream multimodal models show our framework significantly outperforms the previous three common CQA paradigms: instruction-following, OCRenhanced, and chain-of-thought, highlighting the importance of flexible chart understanding for real-world CQA.These findings suggest new directions for developing more robust chart reasoning in future research.
Jingxuan Wei, Junnan Zhu, Haoyanni, Bihui Yu
EMNLP7
2025 ChiImpAVE: An Open-Source Benchmark for Chinese Implicit Attribute Value Extraction
Bihui Yu, Huiyang Shi, Linzhuang Sun, Jingxuan Wei
ICIC (18)1
2025 Beyond Relevance: Utility-Driven Retrieval for Visual Document Question Answering
Bihui Yu, Zhuoya Yao, Huiyang Shi, Liping Bu, Linzhuang Sun, Jingxuan Wei
ICIC (16)1
2025 EEGTCT: Electroencephalogram-Based Chinese Text Decoding
Bihui Yu, Jingxuan Wei, Linzhuang Sun, Liping Bu
ICIC (21)2
2025 SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches
abstract
Hand-drawn sketches are a natural and efficient medium for capturing and conveying ideas. Despite significant advancements in controllable natural image generation, translating freehand sketches into structured, machine-readable diagrams remains a labor-intensive and predominantly manual task. The primary challenge stems from the inherent ambiguity of sketches, which lack the structural constraints and semantic precision required for automated diagram generation. To address this challenge, we introduce SketchAgent, a multi-agent system designed to automate the transformation of hand-drawn sketches into structured diagrams. SketchAgent integrates sketch recognition, symbolic reasoning, and iterative validation to produce semantically coherent and structurally accurate diagrams, significantly reducing the need for manual effort. To evaluate the effectiveness of our approach, we propose the Sketch2Diagram Benchmark, a comprehensive dataset and evaluation framework encompassing eight diverse diagram categories, such as flowcharts, directed graphs, and model architectures. The dataset comprises over 6,000 high-quality examples with token-level annotations, standardized preprocessing, and rigorous quality control. By streamlining the diagram generation process, SketchAgent holds great promise for applications in design, education, and engineering, while offering a significant step toward bridging the gap between intuitive sketching and machine-readable diagram generation.
Cheng Tan 0012, Jingxuan Wei, Zhangyang Gao, Siyuan Li 0002, Bihui Yu, Ruifeng Guo, Stan Z. Li
IJCAI7
2025 ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific Inference
abstract
Understanding how scientific ideas evolve requires more than summarizing individual papers-it demands structured, cross-document reasoning over thematically related research. In this work, we formalize multi-document scientific inference, a new task that extracts and aligns motivation, methodology, and experimental results across related papers to reconstruct research development chains. This task introduces key challenges, including temporally aligning loosely structured methods and standardizing heterogeneous experimental tables. We present ResearchPulse, an agent-based framework that integrates instruction planning, scientific content extraction, and structured visualization. It consists of three coordinated agents: a Plan Agent for task decomposition, a Mmap-Agent that constructs motivation-method mind maps, and a Lchart-Agent that synthesizes experimental line charts. To support this task, we introduce ResearchPulse-Bench, a citation-aware benchmark of annotated paper clusters. Experiments show that our system, despite using 7B-scale agents, consistently outperforms strong baselines like GPT-4o in semantic alignment, structural consistency, and visual fidelity. The dataset are available in https://huggingface.co/datasets/ResearchPulse/ResearchPulse-Bench
Jingxuan Wei, Zhuoya Yao, Bihui Yu, Siyuan Li 0002, Cheng Tan 0012
ACM Multimedia6
2025 Brain-inspired computing based on deep learning for human-computer interaction: A review
Bihui Yu, Jingxuan Wei, Linzhuang Sun, Liping Bu
Neurocomputing1
2024 Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-consistency Training
Cheng Tan 0012, Jingxuan Wei, Zhangyang Gao, Linzhuang Sun, Siyuan Li 0002, Ruifeng Guo, Bihui Yu, Stan Z. Li
ECCV (40)7
2024 Sentence-Level or Token-Level? A Comprehensive Study on Knowledge Distillation
Jingxuan Wei, Linzhuang Sun, Yichong Leng, Xu Tan 0003, Bihui Yu, Ruifeng Guo
IJCAI5
2024 Interpretable and Generalizable Spatiotemporal Predictive Learning with Disentangled Consistency
Jingxuan Wei, Cheng Tan 0012, Zhangyang Gao, Linzhuang Sun, Bihui Yu, Ruifeng Guo, Stan Z. Li
ECML/PKDD (3)5
2024 SAM-Wav2lip++: Enhancing Behavioral Realism in Synthetic Agents Through Audio-Driven Speech and Action Refinement
abstract
Digital human generation is a forward-looking field in technology. Despite significant progress in the generation of speaking facial videos, many challenges remain unaddressed. Issues such as unnatural head movements, distorted expressions, artifacts in generated videos, and uncoordinated limb movements persist. Most current efforts are focused on specific individuals, with enhancements often limited to head movements without further advancing the overall behavioral actions of digital humans. In this context, we introduce a new dataset, CFMD, and a novel model, SAM-Wav2lip++, capable of generating consistent, audio-synchronized lip and behavior action videos from a single reference image of any identity. This work features three main innovative components: (1) a contrastive lip-sync discriminator for precise lip synchronization, (2) a generator for the synthesis of sound-action consistency, and (3) the SAM module for facial refinement operations. Through extensive experiments and user studies, our results demonstrate that our model can synthesize digital human videos of impressively high perceptual quality that accurately sync lip movements and behavioral actions with the input audio, substantially outperforming the state-of-the-art baselines evaluations.
Bihui Yu, Huiyang Shi, Guiyong Chang, Jingxuan Wei, Linzhuang Sun, Songtao Tian, Liping Bu
SMC1
2024 Faster and More Efficient Subject Image Generation for Text-to-Image Diffusion Models
abstract
In recent years, there has been significant progress in text-to-image generation models. However, text struggles to accurately describe abstract concepts like shapes and sizes. Some methods have been proposed to enhance text prompt by incorporating image prompts. While they have shown effective improvements, they either require substantial fine-tuning costs or struggle to effectively integrate text and image information. In our study, we delve into the issue of the difficulty in integrating text and image information in decoupled cross-attention and conduct visual analysis. We identify the presence of background-related tokens in image features as a key factor affecting text fidelity. To address this issue, we develop an algorithm to filter out these tokens. Additionally, we observe differences in the attention of Unet layers to text prompts and image prompts. Based on this finding, we optimize the flow of image information to reduce interference with text information. In summary, we introduce a new topic-customized method that requires no repeated training. It trains a plug-and-play image prompt adapter with only 417M parameters, lightweight yet powerful, surpassing existing models in both text and image consistency. Our code and pre-trained checkpoints will be available at https://github.com/YZBPXX/DDCA.
Bihui Yu, Zhengbing Yao, Jingxuan Wei, Linzhuang Sun, Liping Bu
SMC1
2024 Enhancing human-like multimodal reasoning: a new challenging dataset and comprehensive framework
Jingxuan Wei, Cheng Tan 0012, Zhangyang Gao, Linzhuang Sun, Siyuan Li 0002, Bihui Yu, Ruifeng Guo, Stan Z. Li
Neural Comput. Appl.6
2023 TED-CS: Textual Enhanced Sensitive Video Detection with Common Sense Knowledge
Bihui Yu, Linzhuang Sun, Jingxuan Wei, Shuyue Tan, Yiman Zhao, Liping Bu
ADMA (2)1