Yingfei Liu

dblp:13/5577 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
abstract
Robotic manipulation requires precise spatial understanding to interact with objects in the real world. Point-based methods suffer from sparse sampling, leading to the loss of fine-grained semantics. Image-based methods typically feed RGB and depth into 2D backbones pre-trained on 3D auxiliary tasks, but their entangled semantics and geometry are sensitive to inherent depth noise in real-world that disrupts semantic understanding. Moreover, these methods focus on high-level geometry while overlooking low-level spatial cues essential for precise interaction. We propose SpatialActor, a disentangled framework for robust robotic manipulation that explicitly decouples semantics and geometry. The Semantic-guided Geometric Module adaptively fuses two complementary geometry from noisy depth and semantic-guided expert priors. Also, a Spatial Transformer leverages low-level spatial cues for accurate 2D-3D mapping and enables interaction among spatial features. We evaluate SpatialActor on multiple simulation and real-world scenarios across 50+ tasks. It achieves state-of-the-art performance with 87.4% on RLBench and improves by 13.9% to 19.4% under varying noisy conditions, showing strong robustness. Moreover, it significantly enhances few-shot generalization to new tasks and maintains robustness under various spatial perturbations.
Yingfei Liu, Tiancai Wang, Haoqiang Fan, Xiangyu Zhang 0005, Gao Huang 0001
AAAI3
2026 Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
abstract
The field of autonomous driving increasingly demands high-quality annotated video training data. In this paper, we propose Panacea+, a powerful and universally applicable framework for generating video data in driving scenes. Built upon the foundation of our previous work, Panacea, Panacea+ adopts a multi-view appearance noise prior mechanism and a super-resolution module for enhanced consistency and increased resolution. Extensive experiments show that the generated video samples from Panacea+ greatly benefit a wide range of tasks on different datasets, including 3D object tracking, 3D object detection, and lane detection tasks on the nuScenes and Argoverse 2 dataset. These results strongly prove Panacea+ to be a valuable data generation framework for autonomous driving.
Yuqing Wen, Yingfei Liu, Binyuan Huang, Fan Jia 0006, Chi Zhang 0026, Tiancai Wang, Xiaoyan Sun 0001, Xiangyu Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.3
2025 SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
abstract
Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production in a way that could continuously improve autonomous driving applications. We investigate the impact of scaling up the quantity of generative data on the performance of downstream perception models and find that enhancing data diversity plays a crucial role in effectively scaling generative data production. Therefore, we have developed a novel model equipped with a subject control mechanism, which allows the generative model to leverage diverse external data sources for producing varied and useful data. Extensive evaluations confirm SubjectDrive's efficacy in generating scalable autonomous driving training data, marking a significant step toward revolutionizing data production methods in this field.
Binyuan Huang, Yuqing Wen, Yaosi Hu, Yingfei Liu, Fan Jia 0006, Weixin Mao, Tiancai Wang, Chi Zhang 0026, Chang Wen Chen, Zhenzhong Chen 0001, Xiangyu Zhang 0005
AAAI5
2025 Language Prompt for Autonomous Driving
abstract
A new trend in the computer vision community is to capture objects of interest following flexible human command represented by a natural language prompt. However, the progress of using language prompts in driving scenarios is stuck in a bottleneck due to the scarcity of paired prompt-instance data. To address this challenge, we propose the first object-centric language prompt set for driving scenes within 3D, multi-view, and multi-frame space, named NuPrompt. It expands nuScenes dataset by constructing a total of 40,147 language descriptions, each referring to an average of 7.4 object tracklets. Based on the object-text pairs from the new benchmark, we formulate a novel prompt-based driving task, \ie, employing a language prompt to predict the described object trajectory across views and frames. Furthermore, we provide a simple end-to-end baseline model based on Transformer, named PromptTrack. Experiments show that our PromptTrack achieves impressive performance on NuPrompt. We hope this work can provide some new insights for the self-driving community.
Dongming Wu 0005, Wencheng Han, Yingfei Liu, Tiancai Wang, Cheng-Zhong Xu 0001, Xiangyu Zhang 0005, Jianbing Shen
AAAI3
2025 RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General Grasping
abstract
General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-world effectiveness. To address this limitation, we build a large-scale grasping-oriented affordance segmentation benchmark with human-like instructions, named RAGNet. It contains 273k images, 180 categories, and 26k reasoning instructions. The images cover diverse embodied data domains, such as wild, robot, ego-centric, and even simulation data. They are carefully annotated with an affordance map, while the difficulty of language instructions is largely increased by removing their category name and only providing functional descriptions. Furthermore, we propose a comprehensive affordance-based grasping framework, named AffordanceNet, which consists of a VLM pre-trained on our massive affordance data and a grasping network that conditions an affordance map to grasp the target. Extensive experiments on affordance segmentation benchmarks and real-robot manipulation tasks show that our model has a powerful open-world generalization ability. Our data and code is available at https://github.com/wudongming97/AffordanceNet.
Dongming Wu 0005, Yanping Fu, Saike Huang, Yingfei Liu, Fan Jia 0006, Nian Liu 0002, Tiancai Wang, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jianbing Shen
ICCV4
2025 Glad: A Streaming Scene Generator for Autonomous Driving
abstract
The generation and simulation of diverse real-world scenes have significant application value in the field of autonomous driving, especially for the corner cases. Recently, researchers have explored employing neural radiance fields or diffusion models to generate novel views or synthetic data under driving scenes. However, these approaches suffer from unseen scenes or restricted video length, thus lacking sufficient adaptability for data generation and simulation. To address these issues, we propose a simple yet effective framework, named Glad, to generate video data in a frame-by-frame style. To ensure the temporal consistency of synthetic video, we introduce a latent variable propagation module, which views the latent features of previous frame as noise prior and injects it into the latent features of current frame. In addition, we design a streaming data sampler to orderly sample the original image in a video clip at continuous iterations. Given the reference frame, our Glad can be viewed as a streaming simulator by generating the videos for specific scenes. Extensive experiments are performed on the widely-used nuScenes dataset. Experimental results demonstrate that our proposed Glad achieves promising performance, serving as a strong baseline for online video generation. We will release the source code and models publicly.
Yingfei Liu, Tiancai Wang, Jiale Cao, Xiangyu Zhang 0005
ICLR2
2025 PADriver: Towards Personalized Autonomous Driving
abstract
In this paper, we propose PADriver, a novel closed-loop framework for personalized autonomous driving (PAD). Built upon Multi-modal Large Language Model (MLLM), PADriver takes streaming frames and personalized textual prompts as inputs. It autoaggressively performs scene understanding, danger level estimation and action decision. The predicted danger level reflects the risk of the potential action and provides an explicit reference for the final action, which corresponds to the preset personalized prompt. Moreover, we construct a closed-loop benchmark named PAD-Highway based on Highway-Env simulator to comprehensively evaluate the decision performance under traffic rules. The dataset contains 250 hours videos with high-quality annotation to facilitate the development of PAD behavior analysis. Experimental results on the constructed benchmark show that PADriver outperforms state-of-the-art approaches on different evaluation metrics, and enables various driving modes.
Genghua Kou, Weixin Mao, Yingfei Liu, Osamu Yoshie, Tiancai Wang
IJCNN4
2025 Capacity Enhancement for High-Speed Train Communications With Irregular RIS-Assisted Cell-Free Massive MIMO
abstract
The rapid development of railways worldwide has generated significant attention toward high-speed train (HST) communications. Disruptive technologies for the sixth-generation network, such as cell-free massive multiple-input–multiple-output (mMIMO) and reconfigurable intelligent surface (RIS), are set to substantially enhance the data transmission rate in wireless communication systems. In this article, we leverage the potential benefits of cell-free mMIMO and irregular RIS to achieve the ultimate performance benchmark in HST networks. Specifically, we investigate a cell-free mMIMO system supported by irregular RIS with many low-cost reflecting elements, aiming to improve the weighted sum-rate (WSR) in HST communications. The deployment of irregular RIS allows signals to be transmitted through reflected links with additional degrees of freedom, and thus, a more diverse transmission is facilitated to enhance the received signal power and the system capacity. A joint optimization problem is then formulated to maximize the WSR, while satisfying the constraints on access point power. To address the problem characterized by nonconvexity and high complexity, an alternating optimization algorithm is proposed to iteratively optimize the RIS topology, the transmit beamforming, and the phase shifts. The corresponding subproblems are tackled by exploiting improved tabu search, projected subgradient, and nonmonotone accelerated proximal gradient methods, respectively. Extensive simulations confirm the significant advantages of irregular RIS configurations on the WSR.
Shaowei Li, Yingfei Liu, Yi Liu 0006
IEEE Internet Things J.4
2024 Far3D: Expanding the Horizon for Surround-View 3D Object Detection
abstract
Recently 3D object detection from surround-view images has made notable advancements with its low deployment cost. However, most works have primarily focused on close perception range while leaving long-range detection less explored. Expanding existing methods directly to cover long distances poses challenges such as heavy computation costs and unstable convergence. To address these limitations, this paper proposes a novel sparse query-based framework, dubbed Far3D. By utilizing high-quality 2D object priors, we generate 3D adaptive queries that complement the 3D global queries. To efficiently capture discriminative features across different views and scales for long-range objects, we introduce a perspective-aware aggregation module. Additionally, we propose a range-modulated 3D denoising approach to address query error propagation and mitigate convergence issues in long-range tasks. Significantly, Far3D demonstrates SoTA performance on the challenging Argoverse 2 dataset, covering a wide range of 150 meters, surpassing several LiDAR-based approaches. The code is available at https://github.com/megvii-research/Far3D.
Xiaohui Jiang, Shuailin Li, Yingfei Liu, Fan Jia 0006, Tiancai Wang, Lijin Han, Xiangyu Zhang 0005
AAAI3
2024 Panacea: Panoramic and Controllable Video Generation for Autonomous Driving
abstract
The field of autonomous driving increasingly demands high-quality annotated training data. In this paper, we propose Panacea, an innovative approach to generate panoramic and controllable videos in driving scenarios, capable of yielding an unlimited numbers of diverse, annotated samples pivotal for autonomous driving advancements. Panacea addresses two critical challenges: ‘Consistency’ and ‘Controllability.’ Consistency ensures temporal and cross-view coherence, while Controllability ensures the alignment of generated content with corresponding annotations. Our approach integrates a novel 4D attention and a two-stage generation pipeline to maintain coherence, supplemented by the ControlNet framework for meticulous control by the Bird'View (BEV) layouts. Extensive qualitative and quantitative evaluations of Panacea on the nuScenes dataset prove its effectiveness in generating high-quality multi-view driving-scene videos. This work notably propels the field of autonomous driving by effectively augmenting the training dataset used for advanced BEV perception techniques.
Yuqing Wen, Yingfei Liu, Fan Jia 0006, Chong Luo 0001, Chi Zhang 0026, Tiancai Wang, Xiaoyan Sun 0001, Xiangyu Zhang 0005
CVPR3
2024 Stream Query Denoising for Vectorized HD-Map Construction
Fan Jia 0006, Weixin Mao, Yingfei Liu, Tiancai Wang, Chi Zhang 0026, Xiangyu Zhang 0005, Feng Zhao 0004
ECCV (19)4
2024 TopoMLP: A Simple yet Strong Pipeline for Driving Topology Reasoning
abstract
Topology reasoning aims to comprehensively understand road scenes and present drivable routes in autonomous driving. It requires detecting road centerlines (lane) and traffic elements, further reasoning their topology relationship, \textit{i.e.}, lane-lane topology, and lane-traffic topology. In this work, we first present that the topology score relies heavily on detection performance on lane and traffic elements. Therefore, we introduce a powerful 3D lane detector and an improved 2D traffic element detector to extend the upper limit of topology performance. Further, we propose TopoMLP, a simple yet high-performance pipeline for driving topology reasoning. Based on the impressive detection performance, we develop two simple MLP-based heads for topology generation. TopoMLP achieves state-of-the-art performance on OpenLane-V2 dataset, \textit{i.e.}, 41.2\% OLS with ResNet-50 backbone. It is also the 1st solution for 1st OpenLane Topology in Autonomous Driving Challenge. We hope such simple and strong pipeline can provide some new insights to the community. Code is at https://github.com/wudongming97/TopoMLP.
Dongming Wu 0005, Fan Jia 0006, Yingfei Liu, Tiancai Wang, Jianbing Shen
ICLR4
2023 PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images
abstract
In this paper, we propose PETRv2, a unified framework for 3D perception from multi-view images. Based on PETR [25], PETRv2 explores the effectiveness of temporal modeling, which utilizes the temporal information of previous frames to boost 3D object detection. More specifically, we extend the 3D position embedding (3D PE) in PETR for temporal modeling. The 3D PE achieves the temporal alignment on object position of different frames. To support for multi-task learning (e.g., BEV segmentation and 3D lane detection), PETRv2 provides a simple yet effective solution by introducing task-specific queries, which are initialized under different spaces. PETRv2 achieves state-of-the-art performance on 3D object detection, BEV segmentation and 3D lane detection. Detailed robustness analysis is also conducted on PETR framework. Code is available at https://github.com/megvii-research/PETR.
Yingfei Liu, Fan Jia 0006, Shuailin Li, Aqi Gao, Tiancai Wang, Xiangyu Zhang 0005
ICCV1
2023 Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection
abstract
In this paper, we propose a long-sequence modeling framework, named StreamPETR, for multi-view 3D object detection. Built upon the sparse query design in the PETR series, we systematically develop an object-centric temporal mechanism. The model is performed in an online manner and the long-term historical information is propagated through object queries frame by frame. Besides, we introduce a motion-aware layer normalization to model the movement of the objects. StreamPETR achieves significant performance improvements only with negligible computation cost, compared to the single-frame baseline. On the standard nuScenes benchmark, it is the first online multi-view method that achieves comparable performance (67.6% NDS & 65.3% AMOTA) with lidar-based methods. The lightweight version realizes 45.0% mAP and 31.7 FPS, outperforming the state-of-the-art method (SOLOFusion) by 2.3% mAP and 1.8× faster FPS. Code has been available at https://github.com/exiawsh/StreamPETR.git.
Yingfei Liu, Tiancai Wang, Ying Li 0036, Xiangyu Zhang 0005
ICCV2
2023 Cross Modal Transformer: Towards Fast and Robust 3D Object Detection
abstract
In this paper, we propose a robust 3D detector, named Cross Modal Transformer (CMT), for end-to-end 3D multi-modal detection. Without explicit view transformation, CMT takes the image and point clouds tokens as inputs and directly outputs accurate 3D bounding boxes. The spatial alignment of multi-modal tokens is performed by encoding the 3D points into multi-modal features. The core design of CMT is quite simple while its performance is impressive. It achieves 74.1% NDS (state-of-the-art with single model) on nuScenes test set while maintaining faster inference speed. Moreover, CMT has a strong robustness even if the LiDAR is missing. Code is released at https://github.com/junjie18/CMT.
Yingfei Liu, Jianjian Sun, Fan Jia 0006, Shuailin Li, Tiancai Wang, Xiangyu Zhang 0005
ICCV2
2022 PETR: Position Embedding Transformation for Multi-view 3D Object Detection
Yingfei Liu, Tiancai Wang, Xiangyu Zhang 0005, Jian Sun 0001
ECCV (27)1
2021 SRAF-Net: Shape Robust Anchor-Free Network for Garbage Dumps in Remote Sensing Imagery
abstract
The detection of garbage dumps is of great significance for environmental protection. Recently, deep learning algorithms have brought impressive improvements for regular object detection. Different from conventional objects, garbage dumps are more inconspicuous and irregular and have the problem of blurred boundaries. To solve these problems, we propose a shape robust anchor-free network (SRAF-Net) that consists of feature extraction, multitask detection, and postprocessing. First, our network leverages the context-based deformable (CBD) module to combine context attention and deformable convolution. The contextual information obtained by context attention enables the network to focus on objects with inconspicuous appearance, while the deformable convolution enhances the feature representation. Then, we propose a multitask detection head to regress irregular garbage dumps in a more accurate and efficient way. The anchor-based methods need to define some anchors with a fixed shape. However, our detection method is anchor-free that learns the shapes of objects from training data. The detection head adaptively generates various shapes of bounding boxes with their classification confidences and localization confidences. Weighted by the localization confidences, we merge bounding boxes during postprocessing, which alleviates the blurred boundaries. In addition, we build a new public data set named garbage dumps data set (GDD) to verify the effectiveness of our method. Extensive experiments on GDD indicate that our method surpasses the existing detection methods in terms of speed and accuracy for the garbage dumps detection task.
Xian Sun 0001, Yingfei Liu, Peijin Wang, Wenhui Diao, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 Inversion of Chromophoric Dissolved Organic Matter Using Sparse Regression
abstract
Chromophoric dissolved organic matter (CDOM) retrieval remains to be a challenging task in water color remote sensing research due to its highly spatial and temporal variability. In this paper, we present a novel CDOM retrieval algorithm that takes advantage of the sparse learning, which can simultaneously perform feature selection and parameter estimation. More specifically, by incorporating the band interaction terms into the original spectral matrix and let it be the basis matrix, then the inversion task can be converted to a classical sparse regression problem, namely LASSO, which can be efficiently solved by the coordinate descend algorithm. Experimental results conducted on both simulated and in-situ datasets have demonstrated the efficiency and superiority of the proposed method over some conventional empirical algorithms.
Ruru Deng, Yeheng Liang, Yingfei Liu, Yongming Liu
IGARSS5
2011 Overlapped Handwriting Input on Mobile Phones
abstract
In this paper, we propose an overlapped handwriting input method on handheld devices, which allows users to write continuously without breaks on a single size-restricted writing area. 2 issues have been considered during the implementation of the overlapped input method: previous characters on the background may obstruct the clear viewing of current character and the messy overlapped handwriting is difficult to be segmented and recognized. In our method, a quick segmentation method based on an artificial neural network is used to tackle the first problem and a novel system is implemented to recognize the messy handwriting based on the output of an isolated character recognition engine and a language model. The recognition rate for Chinese characters is about 92.5% for a testing database containing GB2312 Chinese characters and other frequently used symbols. The positive feedbacks from testers have also confirmed the validity of the proposed method.
Yanming Zou, Yingfei Liu, Kongqiao Wang
ICDAR2
2010 Stroke++: a hybrid chinese input method for touch screen mobile phones
abstract
In this paper we present Stroke++, a novel hybrid Chinese input method for touch screen mobile phones that leverages hieroglyphic properties of Chinese characters to enable faster and easier input of Chinese characters on mobile phones. By using a special keypad layout, a friendly user interface and an adaptive radical selection algorithm, we achieved a competitive inputting performance compared with currently prevalent mobile Chinese input methods, while keeping a low entry barrier for Chineseinput novices. An extensive evaluation results show that Stroke++ out-performs the state-of-the-art keystroke-based or handwriting recognition-based Chinese character inputting methods, as far as the input speed and convenience are concerned.
Jianwei Niu 0002, Like Zhu, Qifeng Yan, Yingfei Liu, Kongqiao Wang
Mobile HCI4
2008 Automatic text discovering through stroke-based segmentation and text string combination
abstract
In this paper we present a novel framework of automatic text discovering for content-based multimedia application. For single image, the stroke-based binarization and the coarse-to-fine text extraction will collaborate to generate a clean text image for recognition. For image sequence, multi-frame text enhancement is adopted to increase the text/background contrast, and the recognition results are finally refined by the text string combination algorithm to get more precise semantic information. Two prototype demos have been successfully developed on mobile phones. The experimental results on different platforms show the superior performance of the proposed method.
Lei Xu 0009, Yingfei Liu, Kongqiao Wang
ACM Multimedia2