EDBT 2026 Demo / reviewers in the wild / expert
Penghao Wu
dblp:320/7785
· DBLP profile ↗
17ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0001-6478-1166ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 9 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Video-MMMU: Evaluating Knowledge Acquisition from Multidisciplinary Professional VideosabstractKairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Xiang Yue, Bo Li, Yuanhan Zhang, Ziwei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Xiang Yue, Bo Li 0080, Yuanhan Zhang, Ziwei Liu 0002 |
ACL (1) | 2 |
| 2025 | Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMMabstractLarge multimodal models excel in multimodal tasks but face significant computational challenges due to excessive visual tokens. Unlike token reduction methods that focus on token-level redundancy, we identify and study the computation-level redundancy on vision tokens to ensure no information loss. Our key insight is that vision tokens from the pretrained vision encoder do not necessarily require all the heavy operations (e.g., self-attention, FFNs) in decoder-only LMMs and could be processed more lightly with proper designs. We designed a series of experiments to discover and progressively squeeze out the vision-related computation redundancy. Based on our findings, we propose ProxyV, a novel approach that utilizes proxy vision tokens to alleviate the computational burden on original vision tokens. ProxyV enhances efficiency without compromising performance and can even yield notable performance gains in scenarios with more moderate efficiency improvements. Furthermore, the flexibility of ProxyV is demonstrated through its combination with token reduction methods to boost efficiency further. Penghao Wu, Lewei Lu, Ziwei Liu 0002 |
ICML | 1 |
| 2025 | GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection BehaviorabstractMultimodal Large Language Models (MLLMs) have shown great potential in revolutionizing Graphical User Interface (GUI) automation. However, existing GUI models mostly rely on learning from nearly error-free offline trajectories, thus lacking reflection and error recovery capabilities. To bridge this gap, we propose GUI-Reflection, a novel framework that explicitly integrates self-reflection and error correction capabilities into end-to-end multimodal GUI models throughout dedicated training stages: GUI-specific pre-training, offline supervised fine-tuning (SFT), and online reflection tuning. GUI-reflection enables self-reflection behavior emergence with fully automated data generation and learning processes without requiring any human annotation. Specifically, 1) we first propose scalable data pipelines to automatically construct reflection and error correction data from existing successful trajectories. While existing GUI models mainly focus on grounding and UI understanding ability, we propose the GUI-Reflection Task Suite to learn and evaluate reflection-oriented abilities explicitly. 2) Furthermore, we built a diverse and efficient environment for online training and data collection of GUI models on mobile devices. 3) We also present an iterative online reflection tuning algorithm leveraging the proposed environment, enabling the model to continuously enhance its reflection and error correction abilities. Our framework equips GUI agents with self-reflection and correction capabilities, paving the way for more robust, adaptable, and intelligent GUI automation, with all data, models, environments, and tools to be released publicly. Penghao Wu, Shengnan Ma, Jiaheng Yu, Lewei Lu |
NeurIPS | 1 |
| 2025 | Data-driven spiking neural networks for intelligent fault detection in vehicle lithium-ion battery systems
Penghao Wu, Engang Tian, Hongfeng Tao, Yiyang Chen 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Transfer learning-motivated intelligent fault diagnosis framework for cross-domain knowledge distillation
Penghao Wu, Engang Tian, Hongfeng Tao, Yiyang Chen 0001 |
Neural Networks | 1 |
| 2024 | V*: Guided Visual Search as a Core Mechanism in Multimodal LLMsabstractWhen we look around and perform complex tasks, how we see and selectively process what we see is crucial. How-ever, the lack of this visual search mechanism in current multimodal LLMs (MLLMs) hinders their ability to focus on important visual details, especially when handling high-resolution and visually crowded images. To address this, we introduce V*, an LLM- guided visual search mechanism that employs the world knowledge in LLMs for efficient visual querying. When combined with an MLLM, this mechanism enhances collaborative reasoning, contextual understanding, and precise visual grounding. This integration results in a new MLLM meta-architecture, named Show, sEArch, and TelL (SEAL). We further create V* Bench, a benchmark specifically designed to evaluate MLLMs in their ability to process high-resolution images and focus on visual details. Our study highlights the necessity of incorporating visual search capabilities into multimodal systems. The code is available here. Penghao Wu, Saining Xie |
CVPR | 1 |
| 2024 | Generalized Predictive Model for Autonomous DrivingabstractIn this paper, we introduce the first large-scale video prediction model in the autonomous driving discipline. To eliminate the restriction of high-cost data collection and empower the generalization ability of our model, we ac-quire massive data from the web and pair it with diverse and high-quality text descriptions. The resultant dataset accumulates over 2000 hours of driving videos, spanning areas all over the world with diverse weather conditions and traffic scenarios. Inheriting the merits from recent latent diffusion models, our model, dubbed GenAD, handles the challenging dynamics in driving scenes with novel tem-poral reasoning blocks. We showcase that it can general-ize to various unseen driving datasets in a zero-shot man-ner, surpassing general or driving-specific video prediction counterparts. Furthermore, GenAD can be adapted into an action-conditioned prediction model or a motion planner, holding great potential for real-world driving applications. Jiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen 0008, Tianyu Li 0004, Bo Dai 0002, Kashyap Chitta, Penghao Wu, Ping Luo 0002, Jun Zhang 0106, Andreas Geiger 0001, Yu Qiao 0001, Hongyang Li 0001 |
CVPR | 8 |
| 2024 | Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsabstractWe introduce Cambrian-1, a family of multimodal LLMs (MLLMs) designed with a vision-centric approach. While stronger language models can enhance multimodal capabilities, the design choices for vision components are often insufficiently explored and disconnected from visual representation learning research. This gap hinders accurate sensory grounding in real-world scenarios. Our study uses LLMs and visual instruction tuning as an interface to evaluate various visual representations, offering new insights into different models and architectures—self-supervised, strongly supervised, or combinations thereof—based on experiments with over 15 vision models. We critically examine existing MLLM benchmarks, addressing the difficulties involved in consolidating and interpreting results from various tasks. To further improve visual grounding, we propose spatial vision aggregator (SVA), a dynamic and spatially-aware connector that integrates vision features with LLMs while reducing the number of tokens. Additionally, we discuss the curation of high-quality visual instruction-tuning data from publicly available sources, emphasizing the importance of distribution balancing. Collectively, Cambrian-1 not only achieves state-of-the-art performances but also serves as a comprehensive, open cookbook for instruction-tuned MLLMs. We provide model weights, code, supporting tools, datasets, and detailed instruction-tuning and evaluation recipes. We hope our release will inspire and accelerate advancements in multimodal systems and visual representation learning. Peter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Adithya Iyer, Sai Charitha Akula, Shusheng Yang, Jihan Yang, Manoj Middepogu, Ziteng Wang 0008, Xichen Pan, Rob Fergus, Yann LeCun, Saining Xie |
NeurIPS | 3 |
| 2024 | End-to-End Autonomous Driving: Challenges and FrontiersabstractThe autonomous driving community has witnessed a rapid growth in approaches that embrace an end-to-end algorithm framework, utilizing raw sensor input to generate vehicle motion plans, instead of concentrating on individual tasks such as detection and motion prediction. End-to-end systems, in comparison to modular pipelines, benefit from joint feature optimization for perception and planning. This field has flourished due to the availability of large-scale datasets, closed-loop evaluation, and the increasing need for autonomous driving algorithms to perform effectively in challenging scenarios. In this survey, we provide a comprehensive analysis of more than 270 papers, covering the motivation, roadmap, methodology, challenges, and future trends in end-to-end autonomous driving. We delve into several critical challenges, including multi-modality, interpretability, causal confusion, robustness, and world models, amongst others. Additionally, we discuss current advancements in foundation models and visual pre-training, as well as how to incorporate these techniques within the end-to-end driving framework. Li Chen 0008, Penghao Wu, Kashyap Chitta, Bernhard Jaeger, Andreas Geiger 0001, Hongyang Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Think Twice before Driving: Towards Scalable Decoders for End-to-End Autonomous DrivingabstractEnd-to-end autonomous driving has made impressive progress in recent years. Existing methods usually adopt the decoupled encoder-decoder paradigm, where the encoder extracts hidden features from raw sensor data, and the decoder outputs the ego-vehicle's future trajectories or actions. Under such a paradigm, the encoder does not have access to the intended behavior of the ego agent, leaving the burden of finding out safety-critical regions from the massive receptive field and inferring about future situations to the decoder. Even worse, the decoder is usually composed of several simple multi-layer perceptrons (MLP) or GRUs while the encoder is delicately designed (e.g., a combination of heavy ResNets or Transformer). Such an imbalanced resource-task division hampers the learning process. In this work, we aim to alleviate the aforementioned problem by two principles: (1) fully utilizing the capacity of the encoder; (2) increasing the capacity of the decoder. Concretely, we first predict a coarse-grained future position and action based on the encoder features. Then, conditioned on the position and action, the future scene is imagined to check the ramification if we drive accordingly. We also retrieve the encoder features around the predicted coordinate to obtain fine-grained information about the safety-critical region. Finally, based on the predicted future and the retrieved salient feature, we refine the coarse-grained position and action by predicting its offset from ground-truth. The above refinement module could be stacked in a cascaded fashion, which extends the capacity of the decoder with spatial-temporal prior knowledge about the conditioned future. We conduct experiments on the CARLA simulator and achieve state-of-the-art performance in closed-loop benchmarks. Extensive ablation studies demonstrate the effectiveness of each proposed module. Xiaosong Jia, Penghao Wu, Li Chen 0008, Jiangwei Xie, Conghui He, Junchi Yan, Hongyang Li 0001 |
CVPR | 2 |
| 2023 | Policy Pre-training for Autonomous Driving via Self-supervised Geometric Modeling
Penghao Wu, Li Chen 0008, Hongyang Li 0001, Xiaosong Jia, Junchi Yan, Yu Qiao 0001 |
ICLR | 1 |
| 2023 | HDGT: Heterogeneous Driving Graph Transformer for Multi-Agent Trajectory Prediction via Scene EncodingabstractEncoding a driving scene into vector representations has been an essential task for autonomous driving that can benefit downstream tasks e.g., trajectory prediction. The driving scene often involves heterogeneous elements such as the different types of objects (agents, lanes, traffic signs) and the semantic relations between objects are rich and diverse. Meanwhile, there also exist relativity across elements, which means that the spatial relation is a relative concept and need be encoded in a ego-centric manner instead of in a global coordinate system. Based on these observations, we propose Heterogeneous Driving Graph Transformer (HDGT), a backbone modelling the driving scene as a heterogeneous graph with different types of nodes and edges. For heterogeneous graph construction, we connect different types of nodes according to diverse semantic relations. For spatial relation encoding, the coordinates of the node as well as its in-edges are in the local node-centric coordinate system. For the aggregation module in the graph neural network (GNN), we adopt the transformer structure in a hierarchical way to fit the heterogeneous nature of inputs. Experimental results show that HDGT achieves state-of-the-art performance for the task of trajectory prediction, on INTERACTION Prediction Challenge and Waymo Open Motion Challenge. Xiaosong Jia, Penghao Wu, Li Chen 0008, Yu Liu 0015, Hongyang Li 0001, Junchi Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Inharmonious Region Localization by Magnifying Domain DiscrepancyabstractInharmonious region localization aims to localize the region in a synthetic image which is incompatible with surrounding background. The inharmony issue is mainly attributed to the color and illumination inconsistency produced by image editing techniques. In this work, we tend to transform the input image to another color space to magnify the domain discrepancy between inharmonious region and background, so that the model can identify the inharmonious region more easily. To this end, we present a novel framework consisting of a color mapping module and an inharmonious region localization network, in which the former is equipped with a novel domain discrepancy magnification loss and the latter could be an arbitrary localization network. Extensive experiments on image harmonization dataset show the superiority of our designed framework. Jing Liang 0007, Li Niu 0002, Penghao Wu, Fengjun Guo |
AAAI | 3 |
| 2022 | Inharmonious Region Localization via Recurrent Self-Reasoning
Penghao Wu, Li Niu 0002, Jing Liang 0007, Liqing Zhang 0001 |
BMVC | 1 |
| 2022 | Inharmonious Region Localization with Auxiliary Style Feature
Penghao Wu, Li Niu 0002, Liqing Zhang 0001 |
BMVC | 1 |
| 2022 | ST-P3: End-to-End Vision-Based Autonomous Driving via Spatial-Temporal Feature Learning
Shengchao Hu, Li Chen 0008, Penghao Wu, Hongyang Li 0001, Junchi Yan, Dacheng Tao |
ECCV (38) | 3 |
| 2022 | Trajectory-guided Control Prediction for End-to-end Autonomous Driving: A Simple yet Strong BaselineabstractCurrent end-to-end autonomous driving methods either run a controller based on a planned trajectory or perform control prediction directly, which have spanned two separately studied lines of research. Seeing their potential mutual benefits to each other, this paper takes the initiative to explore the combination of these two well-developed worlds. Specifically, our integrated approach has two branches for trajectory planning and direct control, respectively. The trajectory branch predicts the future trajectory, while the control branch involves a novel multi-step prediction scheme such that the relationship between current actions and future states can be reasoned. The two branches are connected so that the control branch receives corresponding guidance from the trajectory branch at each time step. The outputs from two branches are then fused to achieve complementary advantages. Our results are evaluated in the closed-loop urban driving setting with challenging scenarios using the CARLA simulator. Even with a monocular camera input, the proposed approach ranks first on the official CARLA Leaderboard, outperforming other complex candidates with multiple sensors or fusion mechanisms by a large margin. The sourcecode is publicly available at https://github.com/OpenPerceptionX/TCP Penghao Wu, Xiaosong Jia, Li Chen 0008, Junchi Yan, Hongyang Li 0001, Yu Qiao 0001 |
NeurIPS | 1 |