VLDB 2026 Research / reviewers in the wild / expert
Quanyi Li
dblp:270/7691
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation for Few-Shot Action RecognitionabstractFew-shot Action Recognition (FSAR) aims to recognize novel actions from only a few labeled examples, posing challenges due to limited supervision and complex temporal dynamics. Existing methods often adopt a unified motion modeling strategy for both short- and long-term dynamics, overlooking the need to adapt motion pattern extraction to the specific temporal properties inherent to different timescales. This forces models to hedge against multi-scale relevance through exhaustive searches over temporal tuples, followed by heavy spatio-temporal fusion, which substantially increases parameters and computation and ultimately limits efficiency. To this end, we propose the efficient Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation Network (TCV-STA), which comprises four key components: the Temporal Consistency Module (TCM), the Temporal Variation Module (TVM), the Spatio-Temporal Aggregation attention (STA), and the Shifted Window Temporal Attention (SWTA). The TCM captures stable motion patterns to suppress short-term perturbations and enhance temporal consistency for robust motion representation, while the TVM models dynamic motion patterns to highlight long-term variations that improve inter-class discriminability and facilitate intra-class alignment. Built upon these complementary motion cues, the STA selectively aggregates spatial and temporal representations under the guidance of the learned stable and dynamic motion patterns, avoiding global dense fusion. Finally, to address the limited receptive field and discontinuous modeling caused by frame grouping in TCM and TVM, we adapt a SWTA to capture longer-range temporal dependencies and ensure smooth transitions across subaction segments for few-shot action recognition. Experiments demonstrate that TCV-STA achieves competitive accuracy across four widely-used FSAR benchmarks while reducing parameters by up to 27.9% and computational cost by 21.3%, striking a favorable balance between accuracy and efficiency for deployment in resource-constrained scenarios. Kaiwen Dong, Quanyi Li, Yanjing Sun, Xiao Yun, Yu Zhou 0009, Kévin Riou, Xiaofeng Hou, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Towards Autonomous Micromobility through Scalable Urban SimulationabstractMicromobility, which utilizes lightweight mobile machines moving in urban public spaces - such as delivery robots and electric wheelchairs - emerges as a promising alternative to vehicular mobility. Current micromobility depends mostly on human manual operation (in-person or remote control), which raises safety and efficiency concerns when navigating busy urban environments full of unpredictable obstacles and pedestrians. Assisting humans with AI agents in maneuvering micromobility devices presents a viable solution for enhancing safety and efficiency. In this work, we present a scalable urban simulation solution to advance autonomous micromobility. First, we build URBAN-SIM – a high-performance robot learning platform for large-scale training of embodied agents in interactive urban scenes. URBAN-SIM contains three critical modules: Hierarchical Urban Generation pipeline, Interactive Dynamics Generation strategy, and Asynchronous Scene Sampling scheme, to improve the diversity, realism, and efficiency of robot learning in simulation. Then, we propose URBAN-BENCH – a suite of essential tasks and benchmarks to gauge various capabilities of the AI agents in achieving autonomous micromobility. URBAN-BENCH includes eight tasks based on three core skills of the agents: Urban Locomotion, Urban Navigation, and Urban Traverse. We evaluate four robots with heterogeneous embodiments, such as the wheeled and legged robots, across these tasks. Experiments on diverse terrains and urban structures reveal each robot’s strengths and limitations. Project page: https://metadriverse.github.io/urban-sim/. Wayne Wu, Honglin He, Chaoyuan Zhang, Jack He, Seth Z. Zhao, Quanyi Li, Bolei Zhou |
CVPR | 7 |
| 2025 | Gleam: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor ScenesabstractGeneralizable active mapping in complex unknown environments remains a critical challenge for mobile robots. Existing methods, constrained by insufficient training data and conservative exploration strategies, exhibit limited generalizability across scenes with diverse layouts and complex connectivity. To enable scalable training and reliable evaluation, we introduce GLEAM-Bench, the first large-scale benchmark designed for generalizable active mapping with 1,152 diverse 3D scenes from synthetic and real-scan datasets. Building upon this foundation, we propose GLEAM, a unified generalizable exploration policy for active mapping. Its superior generalizability comes mainly from our semantic representations, long-term navigable goals, and randomized strategies. It significantly outperforms state-of-the-art methods, achieving 66.50% coverage (+9.49%) with efficient trajectories and improved mapping accuracy on 128 unseen complex scenes. Project page: https://xiao-chen.tech/gleam/. Xiao Chen 0016, Quanyi Li, Jiangmiao Pang, Tianfan Xue |
ICCV | 3 |
| 2025 | MetaUrban: An Embodied AI Simulation Platform for Urban MicromobilityabstractPublic urban spaces such as streetscapes and plazas serve residents and accommodate social life in all its vibrant variations. Recent advances in robotics and embodied AI make public urban spaces no longer exclusive to humans. Food delivery bots and electric wheelchairs have started sharing sidewalks with pedestrians, while robot dogs and humanoids have recently emerged in the street. **Micromobility** enabled by AI for short-distance travel in public urban spaces plays a crucial component in future transportation systems. It is essential to ensure the generalizability and safety of AI models used for maneuvering mobile machines. In this work, we present **MetaUrban**, a *compositional* simulation platform for the AI-driven urban micromobility research. MetaUrban can construct an *infinite* number of interactive urban scenes from compositional elements, covering a vast array of ground plans, object placements, pedestrians, vulnerable road users, and other mobile agents' appearances and dynamics. We design point navigation and social navigation tasks as the pilot study using MetaUrban for urban micromobility research and establish various baselines of Reinforcement Learning and Imitation Learning. We conduct extensive evaluation across mobile machines, demonstrating that heterogeneous mechanical structures significantly influence the learning and execution of AI policies. We perform a thorough ablation study, showing that the compositional nature of the simulated environments can substantially improve the generalizability and safety of the trained mobile agents. MetaUrban will be made publicly available to provide research opportunities and foster safe and trustworthy embodied AI and micromobility in cities. The code and data have been released. Wayne Wu, Honglin He, Jack He, Chenda Duan, Zhizheng Liu, Quanyi Li, Bolei Zhou |
ICLR | 7 |
| 2025 | Breaking the Fusion Barrier: An Online Algorithm for Fused Normalization and Linear LayersabstractThe performance of Large Language Model (LLM) inference is critically hindered by memory-bound operations, among which normalization layers are a primary bottleneck, especially during the latency-sensitive decoding phase. While deep learning compilers fuse normalization operations into a single kernel, a fundamental fusion barrier prevents them from merging the ubiquitous Normalization and subsequent Linear layer ($N \& L$) pattern. This barrier, rooted in a core data dependency, forces the execution of two separate kernels, incurring prohibitive kernel launch overhead and costly data round-trips to global memory. In this paper, we break this barrier by introducing FlashFusion, a novel online algorithm that reformulates the$N \& L$pattern to be mathematically equivalent to a single-pass computation. Our key insight is to decompose the computation into a set of independent parallel sums, allowing the normalization statistics to be calculated concurrently with the matrix multiplication, thus eliminating the core data dependency. We co-design a high-performance, hardware-aware GPU kernel that efficiently maps this algorithm to modern architectures, leveraging a tiling strategy to maximize the utilization of Tensor Cores and the memory hierarchy. FlashFusion significantly outperforms state-of-the-art compilers like PyTorch Inductor and TensorRT, achieving speedups of up to$3.03 \times$for the$N \& L$pattern in LLM decoding and effectively eliminating the normalization bottleneck. Hanghang Cao, Shihao Gao, Quanyi Li, Mingjie Xing |
ICPADS | 3 |
| 2024 | GenNBV: Generalizable Next-Best-View Policy for Active 3D ReconstructionabstractWhile recent advances in neural radiance field enable realistic digitization for large-scale scenes, the image-capturing process is still time-consuming and labor-intensive. Previous works attempt to automate this process using the Next-Best-View (NBV) policy for active 3D reconstruction. However, the existing NBV policies heavily rely on handcrafted criteria, limited action space, or perscene optimized representations. These constraints limit their cross-dataset generalizability. To overcome them, we propose GenNBV, an end-to-end generalizable NBV policy. Our policy adopts a reinforcement learning (RL)-based framework and extends typical limited action space to 5D free space. It empowers our agent drone to scan from any viewpoint, and even interact with unseen geometries during training. To boost the cross-dataset generalizability, we also propose a novel multi-source state embedding, including geometric, semantic, and action representations. We establish a benchmark using the Isaac Gym simulator with the Houses3K and OmniObject3D datasets to evaluate this NBV policy. Experiments demonstrate that our policy achieves a 98.26% and 97.12% coverage ratio on unseen building-scale objects from these datasets, respectively, outperforming prior solutions. Xiao Chen 0016, Quanyi Li, Tianfan Xue, Jiangmiao Pang |
CVPR | 2 |
| 2024 | Hybrid Internal Model: Learning Agile Legged Locomotion with Simulated Robot ResponseabstractRobust locomotion control depends on accurate state estimations. However, the sensors of most legged robots can only provide partial and noisy observations, making the estimation particularly challenging, especially for external states like terrain frictions and elevation maps. Inspired by the classical Internal Model Control principle, we consider these external states as disturbances and introduce Hybrid Internal Model (HIM) to estimate them according to the response of the robot. The response, which we refer to as the hybrid internal embedding, contains the robot’s explicit velocity and implicit stability representation, corresponding to two primary goals for locomotion tasks: explicitly tracking velocity and implicitly maintaining stability. We use contrastive learning to optimize the embedding to be close to the robot’s successor state, in which the response is naturally embedded. HIM has several appealing benefits: It only needs the robot’s proprioceptions, i.e., those from joint encoders and IMU as observations. It innovatively maintains consistent observations between simulation reference and reality that avoids information loss in mimicking learning. It exploits batch-level information that is more robust to noises and keeps better sample efficiency. It only requires 1 hour of training on an RTX 4090 to enable a quadruped robot to traverse any terrain under any disturbances. A wealth of real-world experiments demonstrates its agility, even in high-difficulty tasks and cases never occurred during the training process, revealing remarkable open-world generalizability. Junfeng Long, Quanyi Li, Liu Cao, Jiawei Gao 0004, Jiangmiao Pang |
ICLR | 3 |
| 2023 | Guarded Policy Optimization with Imperfect Online Demonstrations
Zhenghai Xue, Zhenghao Peng, Quanyi Li, Bolei Zhou |
ICLR | 3 |
| 2023 | TrafficGen: Learning to Generate Diverse and Realistic Traffic ScenariosabstractDiverse and realistic traffic scenarios are crucial for evaluating the AI safety of autonomous driving systems in simulation. This work introduces a data-driven method called TrafficGen for traffic scenario generation. It learns from the fragmented human driving data collected in the real world and then generates realistic traffic scenarios. TrafficGen is an autoregressive neural generative model with an encoder-decoder architecture. In each autoregressive iteration, it first encodes the current traffic context with the attention mechanism and then decodes a vehicle's initial state followed by generating its long trajectory. We evaluate the trained model in terms of vehicle placement and trajectories, and the experimental result shows our method has substantial improvements over baselines for generating traffic scenarios. After training, TrafficGen can also augment existing traffic scenarios, by adding new vehicles and extending the fragmented trajectories. We further demonstrate that importing the generated scenarios into a simulator as an interactive training environment improves the performance and safety of a driving agent learned from reinforcement learning. Model and data are available at https://metadriverse.github.io/trafficgen. Lan Feng, Quanyi Li, Zhenghao Peng, Shuhan Tan, Bolei Zhou |
ICRA | 2 |
| 2023 | ScenarioNet: Open-Source Platform for Large-Scale Traffic Scenario Simulation and ModelingabstractLarge-scale driving datasets such as Waymo Open Dataset and nuScenes substantially accelerate autonomous driving research, especially for perception tasks such as 3D detection and trajectory forecasting. Since the driving logs in these datasets contain HD maps and detailed object annotations which accurately reflect the real-world complexity of traffic behaviors, we can harvest a massive number of complex traffic scenarios and recreate their digital twins in simulation. Compared to the hand-crafted scenarios often used in existing simulators, data-driven scenarios collected from the real world can facilitate many research opportunities in machine learning and autonomous driving. In this work, we present ScenarioNet, an open-source platform for large-scale traffic scenario modeling and simulation. ScenarioNet defines a unified scenario description format and collects a large-scale repository of real-world traffic scenarios from the heterogeneous data in various driving datasets including Waymo, nuScenes, Lyft L5, and nuPlan datasets. These scenarios can be further replayed and interacted with in multiple views from Bird-Eye-View layout to realistic 3D rendering in MetaDrive simulator. This provides a benchmark for evaluating the safety of autonomous driving stacks in simulation before their real-world deployment. We further demonstrate the strengths of ScenarioNet on large-scale scenario generation, imitation learning, and reinforcement learning in both single-agent and multi-agent settings. Code, demo videos, and website are available at https://github.com/metadriverse/scenarionet Quanyi Li, Zhenghao Peng, Lan Feng, Zhizheng Liu, Chenda Duan, Wenjie Mo 0002, Bolei Zhou |
NeurIPS | 1 |
| 2023 | Learning from Active Human Involvement through Proxy Value PropagationabstractLearning from active human involvement enables the human subject to actively intervene and demonstrate to the AI agent during training. The interaction and corrective feedback from human brings safety and AI alignment to the learning process. In this work, we propose a new reward-free active human involvement method called Proxy Value Propagation for policy optimization. Our key insight is that a proxy value function can be designed to express human intents, wherein state- action pairs in the human demonstration are labeled with high values, while those agents’ actions that are intervened receive low values. Through the TD-learning framework, labeled values of demonstrated state-action pairs are further propagated to other unlabeled data generated from agents’ exploration. The proxy value function thus induces a policy that faithfully emulates human behaviors. Human- in-the-loop experiments show the generality and efficiency of our method. With minimal modification to existing reinforcement learning algorithms, our method can learn to solve continuous and discrete control tasks with various human control devices, including the challenging task of driving in Grand Theft Auto V. Demo video and code are available at: https://metadriverse.github.io/pvp. Zhenghao Peng, Wenjie Mo 0002, Chenda Duan, Quanyi Li, Bolei Zhou |
NeurIPS | 4 |
| 2023 | MetaDrive: Composing Diverse Driving Scenarios for Generalizable Reinforcement LearningabstractDriving safely requires multiple capabilities from human and intelligent agents, such as the generalizability to unseen environments, the safety awareness of the surrounding traffic, and the decision-making in complex multi-agent settings. Despite the great success of Reinforcement Learning (RL), most of the RL research works investigate each capability separately due to the lack of integrated environments. In this work, we develop a new driving simulation platform called MetaDrive to support the research of generalizable reinforcement learning algorithms for machine autonomy. MetaDrive is highly compositional, which can generate an infinite number of diverse driving scenarios from both the procedural generation and the real data importing. Based on MetaDrive, we construct a variety of RL tasks and baselines in both single-agent and multi-agent settings, including benchmarking generalizability across unseen scenes, safe exploration, and learning multi-agent traffic. The generalization experiments conducted on both procedurally generated scenarios and real-world scenarios show that increasing the diversity and the size of the training set leads to the improvement of the RL agent's generalizability. We further evaluate various safe reinforcement learning and multi-agent reinforcement learning algorithms in MetaDrive environments and provide the benchmarks. Source code, documentation, and demo video are available at https://metadriverse.github.io/metadrive. Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, Bolei Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Efficient Learning of Safe Driving Policy via Human-AI Copilot Optimization
Quanyi Li, Zhenghao Peng, Bolei Zhou |
ICLR | 1 |
| 2022 | Human-AI Shared Control via Policy DissectionabstractHuman-AI shared control allows human to interact and collaborate with autonomous agents to accomplish control tasks in complex environments. Previous Reinforcement Learning (RL) methods attempted goal-conditioned designs to achieve human-controllable policies at the cost of redesigning the reward function and training paradigm. Inspired by the neuroscience approach to investigate the motor cortex in primates, we develop a simple yet effective frequency-based approach called Policy Dissection to align the intermediate representation of the learned neural controller with the kinematic attributes of the agent behavior. Without modifying the neural controller or retraining the model, the proposed approach can convert a given RL-trained policy into a human-controllable policy. We evaluate the proposed approach on many RL tasks such as autonomous driving and locomotion. The experiments show that human-AI shared control system achieved by Policy Dissection in driving task can substantially improve the performance and safety in unseen traffic scenes. With human in the inference loop, the locomotion robots also exhibit versatile controllable motion skills even though they are only trained to move forward. Our results suggest the promising direction of implementing human-AI shared autonomy through interpreting the learned representation of the autonomous agents. Code and demo videos are available at https://metadriverse.github.io/policydissect Quanyi Li, Zhenghao Peng, Lan Feng, Bolei Zhou |
NeurIPS | 1 |
| 2021 | Learning to Simulate Self-driven Particles System with Coordinated Policy OptimizationabstractSelf-Driven Particles (SDP) describe a category of multi-agent systems common in everyday life, such as flocking birds and traffic flows. In a SDP system, each agent pursues its own goal and constantly changes its cooperative or competitive behaviors with its nearby agents. Manually designing the controllers for such SDP system is time-consuming, while the resulting emergent behaviors are often not realistic nor generalizable. Thus the realistic simulation of SDP systems remains challenging. Reinforcement learning provides an appealing alternative for automating the development of the controller for SDP. However, previous multi-agent reinforcement learning (MARL) methods define the agents to be teammates or enemies before hand, which fail to capture the essence of SDP where the role of each agent varies to be cooperative or competitive even within one episode. To simulate SDP with MARL, a key challenge is to coordinate agents' behaviors while still maximizing individual objectives. Taking traffic simulation as the testing bed, in this work we develop a novel MARL method called Coordinated Policy Optimization (CoPO), which incorporates social psychology principle to learn neural controller for SDP. Experiments show that the proposed method can achieve superior performance compared to MARL baselines in various metrics. Noticeably the trained vehicles exhibit complex and diverse social behaviors that improve performance and safety of the population as a whole. Demo video and source code are available at: https://decisionforce.github.io/CoPO/ Zhenghao Peng, Quanyi Li, Ka-Ming Hui, Bolei Zhou |
NeurIPS | 2 |
| 2020 | Reinforcement-Learning- and Belief-Learning-Based Double Auction Mechanism for Edge Computing Resource AllocationabstractIn recent years, we have witnessed the compelling application of the Internet of Things (IoT) in our daily life, ranging from daily living to industrial production. On account of the computation and power constraints, the IoT devices have to offload their tasks to the remote cloud services. However, the long-distance transmission poses significant challenges for latency-sensitive businesses, such as autonomous driving and industrial control. As a remedy, mobile edge computing (MEC) is deployed at the edge of the network to reduce the transmission delay. With the MEC joining in, how to allocate the limited computing resource of MEC is a critical problem to guarantee efficient working of the whole IoT system. In this article, we formulate the resource management among MEC and IoT devices as a double auction game. Also, for searching the Nash equilibrium, we introduce the experience-weighted attraction (EWA) algorithm performing behind each participant. With this AI method, auction participants acquire and accumulate experience by observing others' behavior and doing introspection, which accelerates the trading policy's learning process of each agent in such an opaque environment. Some simulation results are presented to evaluate the convergence and correctness of our architecture and algorithm. Quanyi Li, Haipeng Yao, Tianle Mai, Chunxiao Jiang, Yan Zhang 0002 |
IEEE Internet Things J. | 1 |