Yanqin Jiang

dblp:323/4685 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 62% Generative modeling · 27% Autonomous driving · 7%
Computer graphics and multimedia
4 papers
Visual content generation and editing · 80% Computer animation and physical simulation · 20%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
2.332024
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion · NeurIPS 2024
Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video · ICLR 2024
STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians · ECCV (36) 2024
Visual content generation and editing › 3d content generation
4d content generation
2.332024
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion · NeurIPS 2024
STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians · ECCV (36) 2024
SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer · ECCV (13) 2024
Machine learning › Generative modeling
diffusion model
1.022024
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion · NeurIPS 2024
SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer · ECCV (13) 2024
Computer vision › 3D vision
3d reconstruction
0.812024
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion · NeurIPS 2024
Machine learning › Generative modeling › spatio-temporal generative modeling
4d generation
0.812024
Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video · ICLR 2024
Computer vision › 3D vision
neural radiance field
0.812024
Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video · ICLR 2024
Machine learning › Generative modeling › diffusion model
video diffusion model
0.812024
Animate3D: Animating Any 3D Model with Multi-view Video Diffusion · NeurIPS 2024
Visual content generation and editing
3d content creation
0.812024
Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video · ICLR 2024
Computer animation and physical simulation
motion transfer
0.812024
SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer · ECCV (13) 2024
Computer vision › 3D vision
3d object detection
0.712023
PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer · AAAI 2023
Robotics › Autonomous driving › perception › 3d perception
bird's-eye-view perception
0.712023
PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer · AAAI 2023
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
multi-view 3d object detection
0.712023
PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer · AAAI 2023
Computer vision › 3D vision › 3d shape representation
polar representation
0.712023
PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer · AAAI 2023
Machine learning › Deep learning architectures and training › attention mechanism
cross-attention
0.212023
PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer · AAAI 2023
Machine learning › Deep learning architectures and training
transformer
0.212023
PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer · AAAI 2023

Methods — techniques the papers use, named apart from their topics

diffusion model · 3.0video-to-4d generation · 1.5video interpolation · 1.5spatiotemporal attention · 1.5sparse control · 1.5score distillation sampling · 1.5neural radiance field · 1.5gaussian splatting · 1.5polar transformer · 0.7multi-scale representation learning · 0.7
YearPublicationVenuePosition
2025 Two-stream transformer tracking with messengers
Miaobo Qiu, Wenyang Luo, Tongfei Liu, Yanqin Jiang, Jiaming Yan, Weiming Hu 0004, Stephen J. Maybank
Image Vis. Comput.4
2024 SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer
Chaohui Yu, Yanqin Jiang, Chenjie Cao, Fan Wang 0019, Xiang Bai
ECCV (13)3
2024 STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians
Yifei Zeng, Yanqin Jiang, Siyu Zhu 0001, Yuanxun Lu, Youtian Lin, Hao Zhu 0004, Weiming Hu 0004, Xun Cao, Yao Yao 0008
ECCV (36)2
2024 Consistent4D: Consistent 360° Dynamic Object Generation from Monocular Video
abstract
In this paper, we present Consistent4D, a novel approach for generating 4D dynamic objects from uncalibrated monocular videos. Uniquely, we cast the 360-degree dynamic object reconstruction as a 4D generation problem, eliminating the need for tedious multi-view data collection and camera calibration. This is achieved by leveraging the object-level 3D-aware image diffusion model as the primary supervision signal for training dynamic Neural Radiance Fields (DyNeRF). Specifically, we propose a cascade DyNeRF to facilitate stable convergence and temporal continuity under the time-discrete supervision signal. To achieve spatial and temporal consistency of the 4D generation, an interpolation-driven consistency loss is further introduced, which aligns the rendered frames with the interpolated frames from a pre-trained video interpolation model. Extensive experiments show that the proposed Consistent4D significantly outperforms previous 4D reconstruction approaches as well as per-frame 3D generation approaches, opening up new possibilities for 4D dynamic object generation from a single-view uncalibrated video. Project page: https://consistent4d.github.io
Yanqin Jiang, Li Zhang 0040, Weiming Hu 0004, Yao Yao 0008
ICLR1
2024 Animate3D: Animating Any 3D Model with Multi-view Video Diffusion
abstract
Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view attributes, and their results suffer from spatiotemporal inconsistency owing to the inherent ambiguity in the supervision signals. In this work, we present Animate3D, a novel framework for animating any static 3D model. The core idea is two-fold: 1) We propose a novel multi-view video diffusion model (MV-VDM) conditioned on multi-view renderings of the static 3D object, which is trained on our presented large-scale multi-view video dataset (MV-Video). 2) Based on MV-VDM, we introduce a framework combining reconstruction and 4D Score Distillation Sampling (4D-SDS) to leverage the multi-view video diffusion priors for animating 3D objects. Specifically, for MV-VDM, we design a new spatiotemporal attention module to enhance spatial and temporal consistency by integrating 3D and video diffusion models. Additionally, we leverage the static 3D model’s multi-view renderings as conditions to preserve its identity. For animating 3D models, an effective two-stage pipeline is proposed: we first reconstruct coarse motions directly from generated multi-view videos, followed by the introduced 4D-SDS to model fine-level motions. Benefiting from accurate motion learning, we could achieve straightforward mesh animation. Qualitative and quantitative experiments demonstrate that Animate3D significantly outperforms previous approaches. Data, code, and models are open-released.
Yanqin Jiang, Chaohui Yu, Chenjie Cao, Fan Wang 0019, Weiming Hu 0004
NeurIPS1
2023 PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer
abstract
3D object detection in autonomous driving aims to reason “what” and “where” the objects of interest present in a 3D world. Following the conventional wisdom of previous 2D object detection, existing methods often adopt the canonical Cartesian coordinate system with perpendicular axis. However, we conjugate that this does not fit the nature of the ego car’s perspective, as each onboard camera perceives the world in shape of wedge intrinsic to the imaging geometry with radical (non perpendicular) axis. Hence, in this paper we advocate the exploitation of the Polar coordinate system and propose a new Polar Transformer (PolarFormer) for more accurate 3D object detection in the bird’s-eye-view (BEV) taking as input only multi-camera 2D images. Specifically, we design a cross-attention based Polar detection head without restriction to the shape of input structure to deal with irregular Polar grids. For tackling the unconstrained object scale variations along Polar’s distance dimension, we further introduce a multi-scale Polar representation learning strategy. As a result, our model can make best use of the Polar representation rasterized via attending to the corresponding image observation in a sequence-to-sequence fashion subject to the geometric constraints. Thorough experiments on the nuScenes dataset demonstrate that our PolarFormer outperforms significantly state-of-the-art 3D object detection alternatives.
Yanqin Jiang, Li Zhang 0040, Zhenwei Miao, Xiatian Zhu, Weiming Hu 0004, Yu-Gang Jiang 0001
AAAI1
2022 Temporal Point Cloud Fusion With Scene Flow for Robust 3D Object Tracking
abstract
Non-visual range sensors such as Lidar have shown the potential to detect, locate and track objects in complex dynamic scenes thanks to their higher stability in comparison with vision-based sensors like cameras. However, due to the disorder, sparsity, and irregularity of the point cloud, it is much more challenging to take advantage of the temporal information in the dynamic 3D point cloud sequences, as it has been done in the image sequences for improving detection and tracking. In this paper, we propose a novel scene-flow-based point cloud feature fusion module to tackle this challenge, based on which a 3D object tracking framework is also achieved to exploit the temporal motion information. Moreover, we carefully designed several training schemes that contribute to the success of this new module by eliminating the issues of overfitting and long-tailed distribution of object categories. Extensive experiments on the public KITTI 3D object tracking dataset demonstrate the effectiveness of the proposed method by achieving superior results to the baselines. The source code is available athttps://github.com/Tsinghua-OpenICV/SharingVan-OpenPCDet.
Yanding Yang, Kun Jiang 0002, Diange Yang, Yanqin Jiang
IEEE Signal Process. Lett.4