Pengpeng Shao

dblp:236/1614 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 29% Language models and text generation · 20% Efficient and distributed learning · 20%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
1.012026
Event Stream based Human Action Recognition: A High-Definition Benchmark Dataset and Algorithms · Int. J. Comput. Vis. 2026
Computer vision › 3D vision
event-based vision
1.012026
Event Stream based Human Action Recognition: A High-Definition Benchmark Dataset and Algorithms · Int. J. Comput. Vis. 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
Two-Stage Regularization-Based Structured Pruning for LLMs · ACL (1) 2026
Computer vision › Vision and language › vision-language model › multimodal large language model
multimodal large language model reasoning
1.012026
AStar: Boosting Multimodal Reasoning with Automated Structured Thinking · AAAI 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
AStar: Boosting Multimodal Reasoning with Automated Structured Thinking · AAAI 2026
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
1.012026
Two-Stage Regularization-Based Structured Pruning for LLMs · ACL (1) 2026
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models · ACL (1) 2025
Computer vision › Vision and language
multimodal fusion
0.912025
Retain, Blend, and Exchange: A Quality-Aware Spatial-Stereo Fusion Approach for Event Stream Recognition · IEEE Trans. Multim. 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models · ACL (1) 2025
Machine learning › Graph learning
graph neural network
0.312025
Retain, Blend, and Exchange: A Quality-Aware Spatial-Stereo Fusion Approach for Event Stream Recognition · IEEE Trans. Multim. 2025
Information retrieval
retrieval noise
0.312025
Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

empirical evaluation · 1.7benchmark construction · 1.7thought cards · 1.0retrieval of reasoning patterns · 1.0regularization · 1.0event stream processing · 1.0transformer · 0.9quality-aware fusion · 0.9graph neural network · 0.9
YearPublicationVenuePosition
2026 AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
abstract
Multimodal large language models excel across diverse domains but struggle with complex visual reasoning tasks. To enhance their reasoning capabilities, current approaches typically rely on explicit search or post-training techniques. However, search-based methods suffer from computational inefficiency due to extensive solution space exploration, while post-training methods demand substantial data, computational resources, and often exhibit training instability. To address these challenges, we propose **AStar**, a training-free, **A**utomatic **S**tructured **t**hinking paradigm for multimod**a**l **r**easoning. Specifically, we introduce novel "thought cards", a lightweight library of high-level reasoning patterns abstracted from prior samples. For each test problem, AStar adaptively retrieves the optimal thought cards and seamlessly integrates these external explicit guidelines with the model’s internal implicit reasoning capabilities. Compared to previous methods, AStar eliminates computationally expensive explicit search and avoids additional complex post-training processes, enabling a more efficient reasoning approach. Extensive experiments demonstrate that our framework achieves 53.9% accuracy on MathVerse (surpassing GPT-4o's 50.2%) and 32.7% on MathVision (outperforming GPT-4o's 30.4%). Further analysis reveals the remarkable transferability of our method: thought cards generated from mathematical reasoning can also be applied to other reasoning tasks, even benefiting general visual perception and understanding. AStar serves as a plug-and-play test-time inference method, compatible with other post-training techniques, providing an important complement to existing multimodal reasoning approaches.
Mingkuan Feng, Guocheng Zhai, Shuai Zhang 0014, Zheng Lian 0004, Fangrui Lv, Pengpeng Shao, Ruihan Jin, Zhengqi Wen, Jianhua Tao 0001
AAAI7
2026 Two-Stage Regularization-Based Structured Pruning for LLMs
abstract
Mingkuan Feng, Jinyang Wu, Siyuan Liu, Shuai Zhang, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mingkuan Feng, Shuai Zhang 0014, Hongjian Fang, Ruihan Jin, Feihu Che, Pengpeng Shao, Zhengqi Wen, Jianhua Tao 0001
ACL (1)8
2026 Event Stream based Human Action Recognition: A High-Definition Benchmark Dataset and Algorithms
Xiao Wang 0014, Shiao Wang, Pengpeng Shao, Lin Zhu 0012, Bo Jiang 0002, Yonghong Tian 0001
Int. J. Comput. Vis.3
2025 Pandora's Box or Aladdin's Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models
abstract
Retrieval-Augmented Generation (RAG) has emerged as a key method to address hallucinations in large language models (LLMs).While recent research has extended RAG models to complex noisy scenarios, these explorations often confine themselves to limited noise types and presuppose that noise is inherently detrimental to LLMs, potentially deviating from real-world retrieval environments and restricting practical applicability.In this paper, we define seven distinct noise types from a linguistic perspective and establish a Noise RAG Benchmark (NoiserBench), a comprehensive evaluation framework encompassing multiple datasets and reasoning tasks.Through empirical evaluation of eight representative LLMs with diverse architectures and scales, we reveal that these noises can be further categorized into two practical groups: noise that is beneficial to LLMs (aka beneficial noise) and noise that is harmful to LLMs (aka harmful noise).While harmful noise generally impairs performance, beneficial noise may enhance several aspects of model capabilities and overall performance.Our analysis offers insights for developing robust RAG solutions and mitigating hallucinations across diverse retrieval scenarios.Code is available at
Shuai Zhang 0014, Feihu Che, Mingkuan Feng, Pengpeng Shao, Jianhua Tao 0001
ACL (1)5
2025 Retain, Blend, and Exchange: A Quality-Aware Spatial-Stereo Fusion Approach for Event Stream Recognition
abstract
Current event stream-based pattern recognition models typically present the event stream as the point cloud, voxel, image, and the like, and formulate multiple deep neural networks to acquire their features. Although considerable results can be achieved in simple cases, however, the performance of the model might be restricted by monotonous modality expressions, sub-optimal fusion, and readout mechanisms. In this article, we put forward a novel dual-stream framework for event stream-based pattern recognition through differentiated fusion, which is called EFV++. It models two common event representations simultaneously, i.e., event images and event voxels. The spatial and three-dimensional stereo information can be separately learned by making use of Transformer and Graph Neural Network (GNN). We believe the features of each representation still contain both efficient and redundant features and a sub-optimal solution may be obtained if we directly fuse them without differentiation. Thus, we divide each feature into three levels and retain high-quality features, blend medium-quality features, and exchange low-quality features. The enhanced dual features will be provided to the fusion Transformer together with bottleneck features. In addition, we introduce a novel hybrid interaction readout mechanism to enhance the diversity of features as final representations. Comprehensive experiments validate that the framework we have proposed attains cutting-edge performance on a variety of extensively utilized event stream-based classification datasets. Particularly, we have realized a freshly pioneering performance on the Bullying10 k dataset, precisely 90.51%, and this outpaces the runner-up by$+2.21\%$.
Lan Chen 0003, Xiao Wang 0014, Pengpeng Shao, Wei Zhang 0161, Yaowei Wang 0001, Yonghong Tian 0001, Jin Tang 0001
IEEE Trans. Multim.4
2024 Multi-level graph contrastive learning
Pengpeng Shao, Jianhua Tao 0001
Neurocomputing1
2024 Bayesian hypernetwork collaborates with time-difference evolutional network for temporal knowledge prediction
Pengpeng Shao, Jianhua Tao 0001
Neural Networks1
2023 Hierarchical graph attention network for temporal knowledge graph reasoning
Pengpeng Shao, Guanjun Li, Dawei Zhang 0001, Jianhua Tao 0001
Neurocomputing1
2023 Adaptive pseudo-Siamese policy network for temporal knowledge prediction
Pengpeng Shao, Feihu Che, Dawei Zhang 0001, Jianhua Tao 0001
Neural Networks1
2022 Tucker decomposition-based temporal knowledge graph completion
Pengpeng Shao, Dawei Zhang 0001, Guohua Yang, Jianhua Tao 0001, Feihu Che
Knowl. Based Syst.1
2020 Improving iForest for Hydrological Time Series Anomaly Detection
Pengpeng Shao, Yupeng Mao
ICA3PP (3)1
2019 Fast RGB-T Tracking via Cross-Modal Correlation Filters
Sulan Zhai, Pengpeng Shao, Xinyan Liang
Neurocomputing2