VLDB 2026 Research / reviewers in the wild / expert
Haoyang Peng
dblp:303/5776
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Vision and language · 76% Language models and text generation · 24% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › visual question answering
chart question answering |
1.0 | 1 | 2026 | StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Computer vision › Vision and language › vision-language model › multimodal large language model
chart understanding |
1.0 | 1 | 2026 | StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation · ICML 2025 |
Computer vision › Vision and language
multimodal evaluation |
0.9 | 1 | 2025 | RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation · ICML 2025 |
Computer vision › Vision and language › multimodal reasoning
multimodal reasoning benchmark |
0.9 | 1 | 2025 | RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation · ICML 2025 |
Natural language and speech › Language models and text generation › large language model
large language model augmentation |
0.3 | 1 | 2026 | StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Methods — techniques the papers use, named apart from their topics
structured triplet representation · 1.0large language model data augmentation · 1.0model evaluation · 0.9benchmark construction · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StructChart: On the Schema, Metric, and Augmentation for Visual Chart UnderstandingabstractCharts are common in literature across various scientific fields, conveying rich information easily accessible to readers. Current chart-related tasks focus on either chart perception that extracts information from the visual charts, or chart reasoning given the extracted data, e.g. in a tabular form. In this paper, we introduce StructChart, a novel framework that leverages Structured Triplet Representations (STR) to achieve a unified and label-efficient approach to chart perception and reasoning tasks, which is generally applicable to different downstream tasks, beyond the question-answering task as specifically studied in peer works. Specifically, StructChart first reformulates the chart data from the tubular form (linearized CSV) to STR, which can friendlily reduce the task gap between chart perception and reasoning. We then propose a Structuring Chart-oriented Representation Metric (SCRM) to quantitatively evaluate the chart perception task performance. To augment the training, we further explore the potential of Large Language Models (LLMs) to enhance the diversity in both chart visual style and statistical information. Extensive experiments on various chart-related tasks demonstrate the effectiveness and potential of a unified chart perception-reasoning paradigm to push the frontier of chart understanding. Renqiu Xia, Haoyang Peng, Hancheng Ye, Mingsheng Li, Xiangchao Yan, Peng Ye 0006, Botian Shi, Yu Qiao 0001, Junchi Yan, Bo Zhang 0069 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning EvaluationabstractReasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning capabilities required for complex, real-world problemsolving, particularly in multi-disciplinary and multimodal contexts. In this paper, we introduce a graduate-level, multi-disciplinary, EnglishChinese benchmark, dubbed as Reasoning Bench (RBench), for assessing the reasoning capability of both language and multimodal models. RBench spans 1,094 questions across 108 subjects for language model evaluation and 665 questions across 83 subjects for multimodal model testing. These questions are meticulously curated to ensure rigorous difficulty calibration, subject balance, and cross-linguistic alignment, enabling the assessment to be an Olympiad-level multidisciplinary benchmark. We evaluate many models such as o1, GPT-4o, DeepSeek-R1, etc. Experimental results indicate that advanced models perform poorly on complex reasoning, especially multimodal reasoning. Even the top-performing model OpenAI o1 achieves only 53.2% accuracy on our multimodal evaluation. Data and code are made publicly available athttps://evalmodels.github.io/rbench/ Menghao Guo 0001, Yi Zhang 0099, Jiaxi Song, Haoyang Peng, Yi-Xuan Deng, Xinzhi Dong, Kiyohiro Nakayama, Zhengyang Geng, Chen Wang 0049, Bolin Ni, Yongming Rao, Houwen Peng, Han Hu 0001, Gordon Wetzstein, Shi-Min Hu 0001 |
ICML | 5 |
| 2025 | FlipBoost: Strengthening Backdoors in LoRA-Tuned Language Models via Bit-Level InjectionabstractBackdoor attacks pose a significant security threat to large language models (LLMs), allowing adversaries to implant malicious behaviors that are triggered by specific inputs. While existing fine-tuning methods often produce unstable backdoors, their reliability remains limited in real-world scenarios. We propose FlipBoost, a reward-guided bit-level attack targeting LoRA-fine-tuned LLMs to amplify the effectiveness of pre-injected backdoor triggers. FlipBoost identifies and injects a minimal set of high-impact bit flips into LoRA adapter parameters based on target logit gain. Without requiring access to training data or model gradients, our method elevates the trigger activation rate from 10% to 91% using only 15-bit modifications. Experiments on GPT-2 with DailyDialog-style prompts validate the attack’s efficiency and precision. FlipBoost exposes a new vector for post-tuning exploitation in parameter-efficient LLMs and highlights the urgent need for integrity verification and adaptive defense mechanisms. Haoyang Peng, Minghui Xu 0001, Yinhao Xiao |
MASS | 1 |
| 2025 | NP-TCMtarget: a network pharmacology platform for exploring mechanisms of action of traditional Chinese medicineabstractThe biological targets of traditional Chinese medicine (TCM) are the core effectors mediating the interaction between TCM and the human body. Identification of TCM targets is essential to elucidate the chemical basis and mechanisms of TCM for treating diseases. Given the chemical complexity of TCM, both in silico high-throughput compound-target interaction predicting models and biological profile-based methods have been commonly applied for identifying TCM targets based on the structural information of TCM chemical components and biological information, respectively. However, the existing methods lack the integration of TCM chemical and biological information, resulting in difficulty in the systematic discovery of TCM action pathways. To solve this problem, we propose a novel target identification model NP-TCMtarget to explore the TCM target path by combining the overall chemical and biological profiles. First, NP-TCMtarget infers TCM effect targets by calculating associations between herb/disease inducible gene expression profiles and specific gene signatures for 8233 targets. Then, NP-TCMtarget utilizes a constructed binary classification model to predict binding targets of herbal ingredients. Finally, we can distinguish TCM direct and indirect targets by comparing the effect targets and binding targets to establish the action pathways of herbal component-direct target-indirect target by mapping TCM targets in the biological molecular network. We apply NP-TCMtarget to the formula XiaoKeAn to demonstrate the power of revealing the action pathways of herbal formula. We expect that this novel model could provide a systematic framework for exploring the molecular mechanisms of TCM at the target level. NP-TCMtarget is available at http://www.bcxnfz.top/NP-TCMtarget. Aoyi Wang, Haoyang Peng, Yingdong Wang, Caiping Cheng, Jinzhong Zhao, Wuxia Zhang, Peng Li 0004 |
Briefings Bioinform. | 2 |
| 2024 | Multi-View Vision Fusion Network: Can 2D Pre-Trained Model Boost 3D Point Cloud Data-Scarce Learning?abstractPoint cloud based 3D deep model has wide applications in many applications such as autonomous driving, house robot, etc. Inspired by the recent prompt learning in natural language processing, this work proposes a novel Multi-view Vision Fusion Network (MvNet) for few-shot 3D point cloud classification. MvNet investigates the possibility of leveraging the off-the-shelf 2D pre-trained models to achieve the few-shot classification, which can alleviate the over-dependence issue of the existing baseline models towards the large-scale annotated 3D point cloud data. Specifically, MvNet first encodes a 3D point cloud into multi-view image features for a number of different views. Then, a novel multi-view prompt fusion module is developed to fuse information from different views effectively to bridge the gap between 3D point cloud data and 2D pre-trained models. A set of 2D image prompts can then be derived to better describe the suitable prior knowledge for a large-scale pre-trained image model for few-shot 3D point cloud classification. Extensive experiments on ModelNet, ScanObjectNN, and ShapeNet datasets demonstrate that MvNet achieves new state-of-the-art performance for 3D few-shot point cloud image classification. The source code of this work is available at https://github.com/invictus717/MetaTransformer. Haoyang Peng, Baopu Li, Bo Zhang 0069, Xin Chen 0040, Tao Chen 0003, Hongyuan Zhu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |