EDBT 2026 Demo / reviewers in the wild / expert
Jinyu Miao
dblp:257/3982
· DBLP profile ↗
12ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-8558-9173ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video GenerationabstractCorner cases are crucial for training and validating autonomous driving systems, yet collecting them from the real world is often costly and hazardous. Editing objects within captured sensor data offers an effective alternative for generating diverse scenarios, commonly achieved through 3D Gaussian Splatting or image generative models. However, these approaches often suffer from limited visual fidelity or imprecise pose control. To address these issues, we propose G^2Editor, a framework designed for photorealistic and precise object editing in driving videos. Our method leverages a 3D Gaussian representation of the edited object as a dense prior, injected into the denoising process to ensure accurate pose control and spatial consistency. A scene-level 3D bounding box layout is employed to reconstruct occluded areas of non-target objects. Furthermore, to guide the appearance details of the edited object, we incorporate hierarchical fine-grained features as additional conditions during generation. Experiments on the Waymo Open Dataset demonstrate that G^2Editor effectively supports object repositioning, insertion, and deletion within a unified framework, outperforming existing methods in both pose controllability and visual quality, while also benefiting downstream data-driven tasks. Jiusi Li, Jackson Jiang, Jinyu Miao, Miao Long, Tuopu Wen, Peijin Jia, Shengxiang Liu, Chun-lei Yu, Maolin Liu, Yuzhan Cai, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IV | 3 |
| 2026 | MaPLocator: Map-Prior Enhanced End-to-End Vehicle Localization with Coarse-to-Fine Pose Refinement
Muxi Tang, Jinyu Miao, Rujun Yan, Le Jia, Mengmeng Yang 0001, Diange Yang, Kun Jiang 0002 |
IV | 2 |
| 2026 | DTCCL: Disengagement-Triggerd Contrastive Continual Learning for Autonomous Bus Planners
Yanding Yang, Weitao Zhou, Xiaomin Guo, Junze Wen, Lang Ding, Zheng Fu, Jinyu Miao, Kun Jiang 0002, Diange Yang |
IV | 9 |
| 2026 | Temporal Range-Point-Voxel Fusion for Unified BEV Scene Perception and Motion PredictionabstractLiDAR-based bird’s-eye-view (BEV) perception has emerged as an appealing approach for practical autonomous driving applications due to its direct leveraging of precise 3D structures and delivering efficient performance. This paradigm aims to jointly determine the semantics and motion states of various traffic participants on BEV grids. However, most existing LiDAR-based BEV perception methods primarily focus on motion prediction, leading to inferior semantic performance. To address this limitation, we propose a novel multi-frame, multi-view, and multi-task unified framework in this work, which enhances scene perception for both improved BEV semantic segmentation and comparative motion prediction performances. Our framework, named temporal range-point-voxel fusion (T-RPVFusion), leverages a sequence of LiDAR sweeps as input and jointly outputs semantic and motion information on BEV grids. In T-RPVFusion, we first introduce a novel multi-view semantic encoder that extracts high-quality semantic features from each LiDAR sweep. These semantic feature maps are then aggregated into an integrated feature map using the proposed bi-layer spatio-temporal pyramid network. Subsequently, the integrated feature map undergoes processing in both the semantic and motion heads and yields corresponding outputs, respectively. Extensive experiments conducted on Waymo and nuScenes show that our method outperforms previous state-of-the-art (SOTA) in terms of BEV semantic segmentation, while concurrently demonstrating comparable performance in motion prediction. Notably, our method achieves a significant improvement on BEV semantic segmentation task, attaining a mIOU of 49.5%, surpassing the previous SOTA with a great margin of + 12.1% mIOU on Waymo Open Dataset. The code is available athttps://github.com/thuwyl/trpvfusion Yunlong Wang 0009, Kun Jiang 0002, Xinyu Jiao, Jinyu Miao, Yining Shi 0002, Zheng Fu, Mengmeng Yang 0001, Tuopu Wen, Diange Yang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | PriorMotion: Generative Class-Agnostic Motion Prediction with Raster-Vector Motion Field PriorsabstractReliable spatial and motion perception is essential for safe autonomous navigation. Recently, class-agnostic motion prediction on bird's-eye view (BEV) cell grids derived from LiDAR point clouds has gained significant attention. However, existing frameworks typically perform cell classification and motion prediction on a per-pixel basis, neglecting important motion field priors such as rigidity constraints, temporal consistency, and future interactions between agents. These limitations lead to degraded performance, particularly in sparse and distant regions. To address these challenges, we introduce \textbf{PriorMotion}, an innovative generative framework designed for class-agnostic motion prediction that integrates essential motion priors by modeling them as distributions within a structured latent space. Specifically, our method captures structured motion priors using raster-vector representations and employs a variational autoencoder with distinct dynamic and static components to learn future motion distributions in the latent space. Experiments on the nuScenes dataset demonstrate that \textbf{PriorMotion} outperforms state-of-the-art methods across both traditional metrics and our newly proposed evaluation criteria. Notably, we achieve improvements of approximately 15.24\% in accuracy for fast-moving objects, an 3.59\% increase in generalization, a reduction of 0.0163 in motion stability, and a 31.52\% reduction in prediction errors in distant regions. Further validation on FMCW LiDAR sensors confirms the robustness of our approach. Kangan Qian, Jinyu Miao, Xinyu Jiao, Ziang Luo, Zheng Fu, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Diange Yang |
ICCV | 2 |
| 2025 | Efficient End-to-end Visual Localization for Autonomous Driving with Decoupled BEV Neural MatchingabstractAccurate localization plays an important role in high-level autonomous driving systems. Conventional map matching-based localization methods solve the poses by explicitly matching map elements with sensor observations, generally sensitive to perception noise, therefore requiring costly hyperparameter tuning. In this paper, we propose an end-to-end localization neural network which directly estimates vehicle poses from surrounding images, without explicitly matching perception results with HD maps. To ensure efficiency and interpretability, a decoupled BEV neural matching-based pose solver is proposed, which estimates poses in a differentiable sampling-based matching module. Moreover, the sampling space is hugely reduced by decoupling the feature representation affected by each DoF of poses. The experimental results demonstrate that the proposed network is capable of performing decimeter level localization with mean absolute errors of 0.19m, 0.13m and 0.39° in longitudinal, lateral position and yaw angle while exhibiting a 68.8% reduction in inference memory usage. Jinyu Miao, Tuopu Wen, Ziang Luo, Kangan Qian, Zheng Fu, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Jin Huang 0002, Diange Yang |
IROS | 1 |
| 2025 | LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion PredictionabstractAccurate spatial and motion understanding is critical for autonomous driving systems. While object-level perception models excel in structured environments, they struggle with open-set categories and often lack precise geometric representation. Occupancy-based, class-agnostic methods offer better scene expressiveness but typically ignore inter-agent interactions and fail to ensure physical consistency in motion predictions, limiting their reliability in complex traffic scenarios. In this paper, we propose LEGO-Motion, a novel class-agnostic motion prediction framework that bridges the gap between instance-level reasoning and occupancy-based modeling. Unlike conventional grid-based methods that treat each cell independently, LEGO-Motion introduces two key components: (1) the Interaction-Augmented Instance Encoder (IaIE), which models interactions among dynamic agents via cross-attention, and (2) the Instance-Enhanced BEV Encoder (IeBE), which improves motion consistency across instances through multi-stage feature fusion. These components enable our model to learn semantically coherent and physically plausible motion fields. Extensive experiments on the nuScenes dataset show that LEGO-Motion achieves a around 6% improvement in motion prediction accuracy over the previous state-of-the-art, while maintaining real-time inference at 21ms. Moreover, our method demonstrates strong generalization on a proprietary FMCW LiDAR benchmark. These results validate LEGO-Motion's effectiveness in capturing both global scene structure and fine-grained motion dynamics, making it a promising foundation for next-generation perception systems. Kangan Qian, Jinyu Miao, Ziang Luo, Zheng Fu, Jinchen Li, Yining Shi 0002, Yunlong Wang 0009, Kun Jiang 0002, Mengmeng Yang 0001, Diange Yang |
IROS | 2 |
| 2025 | EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Drivingabstract3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computationally heavy and label-hungry. In terms of model complexity, OccNets are commonly composed of heavy Conv3D modules or transformers at the voxel level. Moreover, OccNets are supervised with expensive large-scale dense voxel labels. Model and label inefficiencies, caused by excessive network parameters and label annotation requirements, severely hinder the onboard deployment of OccNets. This paper proposes an EFFicient Occupancy learning framework, EFFOcc, that targets minimal network complexity and label requirements while achieving state-of-the-art accuracy. We first propose an efficient fusion-based OccNet that only uses simple 2D operators and improves accuracy to the state-of-the-art on three large-scale benchmarks: Occ3D-nuScenes, Occ3D-Waymo, and OpenOccupancy-nuScenes. On the Occ3D-nuScenes benchmark, the fusion-based model with ResNet-18 as the image backbone has 21.35M parameters and achieves 51.49 in terms of mean Intersection over Union (mIoU). Furthermore, we propose a multi-stage occupancy-oriented distillation to efficiently transfer knowledge to vision-only OccNet. Extensive experiments on occupancy benchmarks show state-of-the-art precision for both fusion-based and vision-based OccNets. For the demonstration of learning with limited labels, we achieve 94.38% of the performance (mIoU = 28.38) of a 100% labeled vision OccNet (mIoU = 30.07) using the same OccNet trained with only 40% labeled sequences and distillation from the fusion-based OccNet. Code is available at https://github.com/synsin0/EFFOcc. Yining Shi 0002, Kun Jiang 0002, Jinyu Miao, Ke Wang 0021, Kangan Qian, Yunlong Wang 0009, Jiusi Li, Tuopu Wen, Mengmeng Yang 0001, Yiliang Xu, Diange Yang |
IROS | 3 |
| 2023 | ALIKE: Accurate and Lightweight Keypoint Detection and Descriptor ExtractionabstractExisting methods detect the keypoints in a non-differentiable way, therefore they can not directly optimize the position of keypoints through back-propagation. To address this issue, we present a partially differentiable keypoint detection module, which outputs accurate sub-pixel keypoints. The reprojection loss is then proposed to directly optimize these sub-pixel keypoints, and the dispersity peak loss is presented for accurate keypoints regularization. We also extract the descriptors in a sub-pixel way, and they are trained with the stable neural reprojection error loss. Moreover, a lightweight network is designed for keypoint detection and descriptor extraction, which can run at 95 frames per second for 640x480 images on a commercial GPU. On homography estimation, camera pose estimation, and visual (re-)localization tasks, the proposed method achieves equivalent performance with the state-of-the-art approaches, while greatly reduces the inference time. Xiaoming Zhao 0003, Xingming Wu, Jinyu Miao, Weihai Chen, Peter C. Y. Chen, Zhengguo Li |
IEEE Trans. Multim. | 3 |
| 2022 | Discriminative and semantic feature selection for place recognition towards dynamic environments
Jinyu Miao, Xingming Wu, Haosong Yue, Zhong Liu 0005, Weihai Chen |
Pattern Recognit. Lett. | 2 |
| 2021 | RaP-Net: A Region-wise and Point-wise Weighting Network to Extract Robust Features for Indoor LocalizationabstractFeature extraction plays an important role in visual localization. Unreliable features on dynamic objects or repetitive regions will interfere with feature matching and challenge indoor localization greatly. To address the problem, we propose a novel network, RaP-Net, to simultaneously predict region-wise invariability and point-wise reliability, and then extract features by considering both of them. We also introduce a new dataset, named OpenLORIS-Location, to train the proposed network. The dataset contains 1553 images from 93 indoor locations. Various appearance changes between images of the same location are included and can help the model to learn the invariability in typical indoor scenes. Experimental results show that the proposed RaP-Net trained with OpenLORIS-Location dataset achieves excellent performance in the feature matching task and significantly outperforms state-of-the-arts feature algorithms in indoor localization. The RaPNet code and dataset are available at https://github.com/ivipsourcecode/RaP-Net. Dongjiang Li, Jinyu Miao, Xuesong Shi, Qiwei Long, Tianyu Cai, Hongfei Yu, Wei Yang 0029, Haosong Yue, Qi Wei 0001, Fei Qiao |
IROS | 2 |
| 2019 | Robust Loop Closure Detection based on Bag of SuperPoints and Graph VerificationabstractLoop closure detection (LCD) is a crucial technique for robots, which can correct accumulated localization errors after long time explorations. In this paper, we propose a robust LCD algorithm based on Bag of SuperPoints and graph verification. The system first extracts interest points and feature descriptors using the SuperPoint neural network. Then a visual vocabulary is trained in an incremental and self-supervised manner considering the relations between consecutive training images. Finally, a topological graph is constructed using matched feature points to verify candidate loop closures obtained by a Bag-of-Words (BoW) framework. Comparative experiments with state-of-the-art LCD algorithms on several typical datasets have been carried out. The results demonstrate that our proposed graph verification method can significantly improve the accuracy of image matching and the overall LCD approach outperforms existing methods. Haosong Yue, Jinyu Miao, Weihai Chen, Changyun Wen |
IROS | 2 |