Fan Jia 0006

dblp:141/7595-6 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-0252-7207ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
abstract
The field of autonomous driving increasingly demands high-quality annotated video training data. In this paper, we propose Panacea+, a powerful and universally applicable framework for generating video data in driving scenes. Built upon the foundation of our previous work, Panacea, Panacea+ adopts a multi-view appearance noise prior mechanism and a super-resolution module for enhanced consistency and increased resolution. Extensive experiments show that the generated video samples from Panacea+ greatly benefit a wide range of tasks on different datasets, including 3D object tracking, 3D object detection, and lane detection tasks on the nuScenes and Argoverse 2 dataset. These results strongly prove Panacea+ to be a valuable data generation framework for autonomous driving.
Yuqing Wen, Yingfei Liu, Binyuan Huang, Fan Jia 0006, Chi Zhang 0026, Tiancai Wang, Xiaoyan Sun 0001, Xiangyu Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.5
2025 SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control
abstract
Autonomous driving progress relies on large-scale annotated datasets. In this work, we explore the potential of generative models to produce vast quantities of freely-labeled data for autonomous driving applications and present SubjectDrive, the first model proven to scale generative data production in a way that could continuously improve autonomous driving applications. We investigate the impact of scaling up the quantity of generative data on the performance of downstream perception models and find that enhancing data diversity plays a crucial role in effectively scaling generative data production. Therefore, we have developed a novel model equipped with a subject control mechanism, which allows the generative model to leverage diverse external data sources for producing varied and useful data. Extensive evaluations confirm SubjectDrive's efficacy in generating scalable autonomous driving training data, marking a significant step toward revolutionizing data production methods in this field.
Binyuan Huang, Yuqing Wen, Yaosi Hu, Yingfei Liu, Fan Jia 0006, Weixin Mao, Tiancai Wang, Chi Zhang 0026, Chang Wen Chen, Zhenzhong Chen 0001, Xiangyu Zhang 0005
AAAI6
2025 RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General Grasping
abstract
General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-world effectiveness. To address this limitation, we build a large-scale grasping-oriented affordance segmentation benchmark with human-like instructions, named RAGNet. It contains 273k images, 180 categories, and 26k reasoning instructions. The images cover diverse embodied data domains, such as wild, robot, ego-centric, and even simulation data. They are carefully annotated with an affordance map, while the difficulty of language instructions is largely increased by removing their category name and only providing functional descriptions. Furthermore, we propose a comprehensive affordance-based grasping framework, named AffordanceNet, which consists of a VLM pre-trained on our massive affordance data and a grasping network that conditions an affordance map to grasp the target. Extensive experiments on affordance segmentation benchmarks and real-robot manipulation tasks show that our model has a powerful open-world generalization ability. Our data and code is available at https://github.com/wudongming97/AffordanceNet.
Dongming Wu 0005, Yanping Fu, Saike Huang, Yingfei Liu, Fan Jia 0006, Nian Liu 0002, Tiancai Wang, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jianbing Shen
ICCV5
2024 Far3D: Expanding the Horizon for Surround-View 3D Object Detection
abstract
Recently 3D object detection from surround-view images has made notable advancements with its low deployment cost. However, most works have primarily focused on close perception range while leaving long-range detection less explored. Expanding existing methods directly to cover long distances poses challenges such as heavy computation costs and unstable convergence. To address these limitations, this paper proposes a novel sparse query-based framework, dubbed Far3D. By utilizing high-quality 2D object priors, we generate 3D adaptive queries that complement the 3D global queries. To efficiently capture discriminative features across different views and scales for long-range objects, we introduce a perspective-aware aggregation module. Additionally, we propose a range-modulated 3D denoising approach to address query error propagation and mitigate convergence issues in long-range tasks. Significantly, Far3D demonstrates SoTA performance on the challenging Argoverse 2 dataset, covering a wide range of 150 meters, surpassing several LiDAR-based approaches. The code is available at https://github.com/megvii-research/Far3D.
Xiaohui Jiang, Shuailin Li, Yingfei Liu, Fan Jia 0006, Tiancai Wang, Lijin Han, Xiangyu Zhang 0005
AAAI5
2024 Panacea: Panoramic and Controllable Video Generation for Autonomous Driving
abstract
The field of autonomous driving increasingly demands high-quality annotated training data. In this paper, we propose Panacea, an innovative approach to generate panoramic and controllable videos in driving scenarios, capable of yielding an unlimited numbers of diverse, annotated samples pivotal for autonomous driving advancements. Panacea addresses two critical challenges: ‘Consistency’ and ‘Controllability.’ Consistency ensures temporal and cross-view coherence, while Controllability ensures the alignment of generated content with corresponding annotations. Our approach integrates a novel 4D attention and a two-stage generation pipeline to maintain coherence, supplemented by the ControlNet framework for meticulous control by the Bird'View (BEV) layouts. Extensive qualitative and quantitative evaluations of Panacea on the nuScenes dataset prove its effectiveness in generating high-quality multi-view driving-scene videos. This work notably propels the field of autonomous driving by effectively augmenting the training dataset used for advanced BEV perception techniques.
Yuqing Wen, Yingfei Liu, Fan Jia 0006, Chong Luo 0001, Chi Zhang 0026, Tiancai Wang, Xiaoyan Sun 0001, Xiangyu Zhang 0005
CVPR4
2024 Stream Query Denoising for Vectorized HD-Map Construction
Fan Jia 0006, Weixin Mao, Yingfei Liu, Tiancai Wang, Chi Zhang 0026, Xiangyu Zhang 0005, Feng Zhao 0004
ECCV (19)2
2024 TopoMLP: A Simple yet Strong Pipeline for Driving Topology Reasoning
abstract
Topology reasoning aims to comprehensively understand road scenes and present drivable routes in autonomous driving. It requires detecting road centerlines (lane) and traffic elements, further reasoning their topology relationship, \textit{i.e.}, lane-lane topology, and lane-traffic topology. In this work, we first present that the topology score relies heavily on detection performance on lane and traffic elements. Therefore, we introduce a powerful 3D lane detector and an improved 2D traffic element detector to extend the upper limit of topology performance. Further, we propose TopoMLP, a simple yet high-performance pipeline for driving topology reasoning. Based on the impressive detection performance, we develop two simple MLP-based heads for topology generation. TopoMLP achieves state-of-the-art performance on OpenLane-V2 dataset, \textit{i.e.}, 41.2\% OLS with ResNet-50 backbone. It is also the 1st solution for 1st OpenLane Topology in Autonomous Driving Challenge. We hope such simple and strong pipeline can provide some new insights to the community. Code is at https://github.com/wudongming97/TopoMLP.
Dongming Wu 0005, Fan Jia 0006, Yingfei Liu, Tiancai Wang, Jianbing Shen
ICLR3
2023 PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images
abstract
In this paper, we propose PETRv2, a unified framework for 3D perception from multi-view images. Based on PETR [25], PETRv2 explores the effectiveness of temporal modeling, which utilizes the temporal information of previous frames to boost 3D object detection. More specifically, we extend the 3D position embedding (3D PE) in PETR for temporal modeling. The 3D PE achieves the temporal alignment on object position of different frames. To support for multi-task learning (e.g., BEV segmentation and 3D lane detection), PETRv2 provides a simple yet effective solution by introducing task-specific queries, which are initialized under different spaces. PETRv2 achieves state-of-the-art performance on 3D object detection, BEV segmentation and 3D lane detection. Detailed robustness analysis is also conducted on PETR framework. Code is available at https://github.com/megvii-research/PETR.
Yingfei Liu, Fan Jia 0006, Shuailin Li, Aqi Gao, Tiancai Wang, Xiangyu Zhang 0005
ICCV3
2023 Cross Modal Transformer: Towards Fast and Robust 3D Object Detection
abstract
In this paper, we propose a robust 3D detector, named Cross Modal Transformer (CMT), for end-to-end 3D multi-modal detection. Without explicit view transformation, CMT takes the image and point clouds tokens as inputs and directly outputs accurate 3D bounding boxes. The spatial alignment of multi-modal tokens is performed by encoding the 3D points into multi-modal features. The core design of CMT is quite simple while its performance is impressive. It achieves 74.1% NDS (state-of-the-art with single model) on nuScenes test set while maintaining faster inference speed. Moreover, CMT has a strong robustness even if the LiDAR is missing. Code is released at https://github.com/junjie18/CMT.
Yingfei Liu, Jianjian Sun, Fan Jia 0006, Shuailin Li, Tiancai Wang, Xiangyu Zhang 0005
ICCV4
2023 Domain-specific feature elimination: multi-source domain adaptation for image classification
Kunhong Wu, Fan Jia 0006, Yahong Han
Frontiers Comput. Sci.2
2021 Adversarial Attack with KD-Tree Searching on Training Set
Xinghua Guo, Fan Jia 0006, Jianqiao An, Yahong Han
ICIG (2)2
2021 Zero Knowledge Adversarial Defense Via Iterative Translation Cycle
abstract
Image classification networks based on deep learning are found to be vulnerable to the carefully designed adversarial examples. The existing approaches are still far from solving this problem. In this paper, we propose iterative translation cycle GAN (ITC-GAN) to jointly optimize generators and discriminators, so as to defense against adversarial examples with zero knowledge of adversarial attacks. We train two generators to form an iterative translation cycle which changes deep features of the input images and then reconstructs them. As the cycle of image translation can be conducted iteratively, the noises of adversarial examples are gradually eliminated. We only use clean images to train the whole ITC-GAN, so our method is not coupled to specific attack methods and specific classifiers. We conduct experiments on MNIST, CIFAR10, and ImageNet50. The experimental results demonstrate that ITC-GAN is more robust and flexible than state-of-the-art methods in different adversarial settings.
Fan Jia 0006, Yahong Han
ICME1
2020 Two-Way Feature-Aligned And Attention-Rectified Adversarial Training
abstract
Adversarial training increases robustness by augmenting training data with adversarial examples. However, vanilla adversarial training may be overfitting to certain adversarial attacks. Small perturbations in images bring in error which is gradually amplified when forwarded through the model so that the error leads to wrong classification. Besides, small perturbations will also distract classifier's attention to significant features that are relevant to the true label. In this paper, we propose a novel two-way feature-aligned and attention-rectified adversarial training (FAAR) to improve adversarial training (AT). FAAR utilizes two-way feature alignment and attention rectification to mitigate the problems mentioned above. FAAR effectively suppresses perturbations in lowlevel, high-level and global features by moving features of perturbed images towards those of clean images with twoway feature alignment. It also leads the model into focusing more on useful features which are correlated with true label through rectifying gradient-weighted attention. Besides, feature alignment activates attention rectification by reducing perturbations in high-level feature. Our proposed method FAAR surpasses other existing AT methods in three aspects. First, it pushes the model to keep invariant when dealing with different adversarial attacks and different magnitude of perturbations. Second, it can be applied to any convolution neural networks. Third, the training process is end-to-end. For experiments, FAAR shows promising defense performance on CIFAR-10 and ImageNet.
Fan Jia 0006, Quanxin Zhang 0001, Yahong Han, Xiaohui Kuang, Yu-an Tan 0001
ICME2