Wenyi Li 0001

dblp:224/5069-1 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-0616-0817ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Robot manipulation · 23% Generative modeling · 20% 3D vision · 15%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › visual question answering
driving question answering
0.912025
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models · NeurIPS 2025
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.912025
PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model · CVPR 2025
Robotics › Robot manipulation › embodied foundation models
vision-language-action model
0.912025
Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.812024
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis · ECCV (32) 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-target domain adaptation
0.812024
Training-Free Model Merging for Multi-target Domain Adaptation · ECCV (47) 2024
Machine learning › Generative modeling › image generation › conditional image synthesis
semantic image synthesis
0.812024
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis · ECCV (32) 2024
Machine learning › Efficient and distributed learning › model merging
training-free model merging
0.812024
Training-Free Model Merging for Multi-target Domain Adaptation · ECCV (47) 2024
Visual content generation and editing
image generation
0.812024
SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis · ECCV (32) 2024
Computer vision › 3D vision › 3d reconstruction › learning-based 3d reconstruction
3d gaussian reconstruction
0.312025
PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model · CVPR 2025

Methods — techniques the papers use, named apart from their topics

spatial-categorical prior · 1.5diffusion · 1.5vision-language-action model · 0.9two-stage training · 0.9trajectory prediction · 0.9drag embedding · 0.93d gaussian splatting · 0.9model merging · 0.8
YearPublicationVenuePosition
2026 FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature Banks
abstract
With the rapid advancements in diffusion models and 3D generation techniques, dynamic 3D content generation has become a crucial research area. However, achieving high-fidelity 4D (dynamic 3D) generation with strong spatial-temporal consistency remains a challenging task. Inspired by recent findings that pretrained diffusion features capture rich correspondences, we propose FB-4D, a novel 4D generation framework that integrates a Feature Bank mechanism to enhance both spatial and temporal consistency in generated frames. In FB-4D, we store features extracted from previous frames and fuse them into the process of generating subsequent frames, ensuring consistent characteristics across both time and multiple views. To ensure a compact representation, the Feature Bank is updated by a proposed dynamic merging mechanism. Leveraging this Feature Bank, we demonstrate for the first time that generating additional reference sequences through multiple autoregressive iterations can continuously improve generation performance. Experimental results show that FB-4D significantly outperforms existing methods in terms of rendering quality, spatial-temporal consistency, and robustness. It surpasses all multi-view generation tuning-free approaches by a large margin and achieves performance on par with training-based methods. Our code and data will be publicly available to support future research.
Huan-ang Gao, Wenyi Li 0001, Haohan Chi, Chenxi Du, Yiqian Liu, Mingju Gao, Guiyu Zhang, Zongzheng Zhang, Li Yi 0001, Hongyang Li 0001, Hao Zhao 0002
WACV3
2025 PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model
abstract
As interest grows in world models that predict future states from current observations and actions, accurately modeling part-level dynamics has become increasingly relevant for various applications. Existing approaches, such as Puppet-Master, rely on fine-tuning large-scale pre-trained video diffusion models, which are impractical for real-world use due to the limitations of 2D video representation and slow processing times. To overcome these challenges, we present PartRM, a novel 4D reconstruction framework that simultaneously models appearance, geometry, and part-level motion from multi-view images of a static object. PartRM builds upon large 3D Gaussian reconstruction models, leveraging their extensive knowledge of appearance and geometry in static objects. To address data scarcity in 4D, we introduce the PartDrag-4D dataset, providing multi-view observations of part-level dynamics across over 20,000 states. We enhance the model’s understanding of interaction conditions with a multi-scale drag embedding module that captures dynamics at varying granularities. To prevent catastrophic forgetting during fine-tuning, we implement a two-stage training process that focuses sequentially on motion and appearance learning. Experimental results show that PartRM establishes a new state-of-the-art in part-level motion learning and can be applied in manipulation tasks in robotics. Our code, data, and models are publicly available to facilitate future research.
Mingju Gao, Yike Pan, Huan-ang Gao, Zongzheng Zhang, Wenyi Li 0001, Hao Dong 0003, Hao Tang 0005, Li Yi 0001, Hao Zhao 0002
CVPR5
2025 Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
abstract
Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution is the Impromptu VLA Dataset: over 80,000 meticulously curated video clips, distilled from over 2M source clips sourced from 8 open-source large-scale datasets. This dataset is built upon our novel taxonomy of four challenging unstructured categories and features rich, planning-oriented question-answering annotations and action trajectories. Crucially, experiments demonstrate that VLAs trained with our dataset achieve substantial performance gains on established benchmarks—improving closed-loop NeuroNCAP scores and collision rates, and reaching near state-of-the-art L2 accuracy in open-loop nuScenes trajectory prediction. Furthermore, our Q&A suite serves as an effective diagnostic, revealing clear VLM improvements in perception, prediction, and planning. Our code, data and models are available at https://github.com/ahydchh/Impromptu-VLA
Haohan Chi, Huan-ang Gao, Kaisen Yang, Yangcheng Yu, Zeda Wang, Wenyi Li 0001, Leichen Wang, Xingtao Hu, Hang Zhao 0021, Hao Zhao 0002
NeurIPS10
2024 SCP-Diff: Spatial-Categorical Joint Prior for Diffusion Based Semantic Image Synthesis
Huan-ang Gao, Mingju Gao, Jiaju Li, Wenyi Li 0001, Rong Zhi, Hao Tang 0005, Hao Zhao 0002
ECCV (32)4
2024 Training-Free Model Merging for Multi-target Domain Adaptation
Wenyi Li 0001, Huan-ang Gao, Mingju Gao, Beiwen Tian, Rong Zhi, Hao Zhao 0002
ECCV (47)1
2024 FairDiff: Fair Segmentation with Point-Image Diffusion
Wenyi Li 0001, Haoran Xu 0003, Guiyu Zhang, Huan-ang Gao, Mingju Gao, Hao Zhao 0002
MICCAI (3)1