VLDB 2026 Research / reviewers in the wild / expert
Wei Tian 0001
dblp:56/3860-1
· DBLP profile ↗
22ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-5085-7219ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AMTFusion: Boosting 3D Object Detection by Adaptive Multi-Modal Temporal Fusion and AugmentationabstractAccurate perception of surrounding 3D objects is indispensable for ensuring the safety of autonomous driving systems. Furthermore, the integration of multi-modal sensors has emerged as a pivotal research area within detection paradigms based on Bird’s Eye View. In this paper, we introduce a sparse query-based 3D object detection framework by adaptive multi-modal temporal fusion of camera and LiDAR sensor data. Initially, we leverage the spatial-aware cross-modal feature aggregation to adaptively fuse instance-level point cloud and image features. We further elaborate a sequential temporal modeling approach, which facilitates an enhanced multi-modal 3D object detection within constant spatial and temporal complexity. To mitigate the challenge posed by limited high-quality training data, we incorporate a fragment-wise multi-modal temporal data augmentation strategy, aimed at bolstering the model’s performance for detecting long-tail distributed objects. Our approach achieves an NDS detection score of 75.5% on the nuScenes test set (without test time augmentation and model ensembling). These results underscore the efficacy of our approach and highlight its superiority compared to prior state-of-the-art methods. We further show that our approach remarkably improves the accuracy of multi-object tracking on nuScenes dataset. Wei Tian 0001, Zhenglin Du, Qiankun Yu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | R2LDM: An Efficient 4D Radar Super-Resolution Framework Leveraging Diffusion ModelabstractWe introduce R2LDM, an innovative approach for generating dense and accurate 4D radar point clouds, guided by corresponding LiDAR point clouds. Instead of utilizing range images or bird’s eye view (BEV) images, we represent both LiDAR and 4D radar point clouds using voxel features, which more effectively capture 3D shape information. Subsequently, we propose the Latent Voxel Diffusion Model (LVDM), which performs the diffusion process in the latent space. Additionally, a novel Latent Point Cloud Reconstruction (LPCR) module is utilized to reconstruct point clouds from high-dimensional latent voxel features. As a result, R2LDM effectively generates LiDAR-like point clouds from paired raw radar data. We evaluate our approach on two different datasets, and the experimental results demonstrate that our model achieves 6- to 10-fold densification of radar point clouds, outperforming state-of-the-art baselines in 4D radar point cloud super-resolution. Furthermore, the enhanced radar point clouds generated by our method significantly improve downstream tasks, achieving up to 31.7% improvement in point cloud registration recall rate and 24.9% improvement in object detection accuracy. Shouyi Lu, Renbo Huang, Minqing Huang, Wei Tian 0001, Guirong Zhuo, Lu Xiong 0001 |
IROS | 6 |
| 2025 | Attention-Based Two-Stage 3D Lane Detection and Topological PredictionabstractThe increasing demand for accurate perception of static road information in autonomous driving systems has drawn significant attention to 3D lane detection and topology prediction. This paper introduces a two-stage 3D lane detection and topology prediction model based on attention mechanism. The proposed model employs 3D Bézier curves to represent lane lines and an adjacency matrix to present the topological relationships, which facilitate the end-to-end learning of detection and topology prediction for 3D lanes. Experimental results on the OpenLaneV2 dataset demonstrate that the proposed method achieves improvements in terms of both 3D lane detection and topology prediction compared to currently outstanding methods. Xiaohan Fu, Wei Tian 0001, Xianwang Yu |
IV | 3 |
| 2025 | Fully Convolutional Neural Network-Based Speech Enhancement for In-Vehicle Environment in 3D PerspectiveabstractVoice interaction is one of the important development directions of intelligent cabins. However, various noise interferences inside and outside the vehicle pose significant challenges to human-vehicle interaction. Speech enhancement technology can extract clear speech signals from mixed speech signals, thereby significantly improving the quality and intelligibility of voice commands, making it a current research hotspot. To address the issue of low intelligibility of voice commands in vehicle environment, we propose a speech enhancement method based on 3D tensor representation. Specifically, through short-time Fourier transform, the real and imaginary parts are concatenated in a new dimension to form a 3D tensor as the input feature. Based on convolutional neural networks, we conducted systematic research, building UNet speech enhancement models based on 1D, 2D, and 3D convolutions respectively, and compared the performance of the proposed 3D complex domain model with that of the time-domain and frequency-domain models. Additionally, we incorporated an attention mechanism and developed a time-frequency domain joint model by leveraging its information filtering and focusing capabilities, thereby significantly enhancing the speech enhancement effect. Finally, we conducted extensive experiments on the Voicebank+Demand dataset. The results show that although the time-frequency domain joint model outperforms the 3D complex domain model on the test set, in the vehicle environment, the 3D complex domain model achieved the highest scores of 3.67, 97.94, and 4.36 respectively in the three metrics of PESQ, STOI, and COVL. Kaikun Pei, Dejian Meng, Wei Tian 0001 |
IV | 4 |
| 2025 | Dual-sampling feature fusion for three-dimensional object detection using four-dimensional radar and camera
Caien Weng, Panpan Tong, Wei Tian 0001, Lu Xiong 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose EstimationabstractLocating 3D objects from a single RGB image via Perspective-n-Point (PnP) is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest interpreting PnP as a differentiable layer, allowing for partial learning of 2D-3D point correspondences by backpropagating the gradients of pose loss. Yet, learning the entire correspondences from scratch is highly challenging, particularly for ambiguous pose solutions, where the globally optimal pose is theoretically non-differentiable w.r.t. the points. In this paper, we propose the EPro-PnP, a probabilistic PnP layer for general end-to-end pose estimation, which outputs a distribution of pose with differentiable probability density on the SE(3) manifold. The 2D-3D coordinates and corresponding weights are treated as intermediate variables learned by minimizing the KL divergence between the predicted and target pose distribution. The underlying principle generalizes previous approaches, and resembles the attention mechanism. EPro-PnP can enhance existing correspondence networks, closing the gap between PnP-based method and the task-specific leaders on the LineMOD 6DoF pose estimation benchmark. Furthermore, EPro-PnP helps to explore new possibilities of network design, as we demonstrate a novel deformable correspondence network with the state-of-the-art pose accuracy on the nuScenes 3D object detection benchmark. Hansheng Chen 0001, Wei Tian 0001, Pichao Wang, Fan Wang 0019, Lu Xiong 0001, Hao Li 0030 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Spectral Scaling-Based Augmentation for Corruption-Robust Image ClassificationabstractImage classifiers often degrade in performance when test images differ significantly from the training distribution due to real-world image corruptions. Frequency-based augmentations can be used to address this issue, but existing methods excel against corruptions caused by noise and blur while struggling with those caused by contrast and fog. To tackle these challenges, we propose a novel image augmentation method grounded in a new perspective of relative spectral differences. This perspective characterizes spectral variations introduced by common corruptions as changes in non-zero frequencies, providing a unified understanding of their effects on image spectra. Building on this insight, the proposed method incorporates two key modules: a random spectral scaling module that captures statistical properties of image spectra and a deep spectral scaling module that adaptively learns spectral adjustments through a neural network. Experiments demonstrate that the proposed method improves overall robustness across various corruptions, with notable gains of 6.3% and 6.4% on contrast and fog, respectively, where existing methods often fall short. Dejian Meng, Wei Tian 0001, Jun Yan 0013 |
IEEE Signal Process. Lett. | 4 |
| 2024 | DiffusionRegPose: Enhancing Multi-Person Pose Estimation Using a Diffusion-Based End-to-End Regression ApproachabstractThis paper presents the DiffusionRegPose, a novel approach to multi-person pose estimation that converts a one-stage, end-to-end keypoint regression model into a diffusion-based sampling process. Existing one-stage deterministic re-gression methods, though efficient, are often prone to missed or false detections in crowded or occluded scenes, due to their inability to reason pose ambiguity. To address these challenges, we handle ambiguous poses in a generative fashion, i.e., sampling from the image-conditioned pose distributions characterized by a diffusion probabilistic model. Specifically, with initial pose tokens extracted from the image, noisy pose candidates are progressively refined by inter-acting with the initial tokens via attention layers. Extensive evaluations on the COCO and CrowdPose datasets show that DiffusionRegPose clearly improves the pose accuracy in crowded scenarios, as evidenced by a notable 4. 0 AP in-crease in the APHmetric on the CrowdPose dataset. This demonstrates the model's potential for robust and precise human pose estimation in real-world applications. Code will be available at https://github.com/cici203IDiffusionRegPose. Dayi Tan, Hansheng Chen 0001, Wei Tian 0001, Lu Xiong 0001 |
CVPR | 3 |
| 2024 | Refinecurvelane: lane detection with B-spline curve in a layer-by-layer refinement manner
Wei Tian 0001, Yuyao Huang 0001, Xianwang Yu |
Multim. Syst. | 1 |
| 2023 | Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstructionabstract3D-aware image synthesis encompasses a variety of tasks, such as scene generation and novel view synthesis from images. Despite numerous task-specific methods, developing a comprehensive model remains challenging. In this paper, we present SSDNeRF, a unified approach that employs an expressive diffusion model to learn a generalizable prior of neural radiance fields (NeRF) from multi-view images of diverse objects. Previous studies have used two-stage approaches that rely on pretrained NeRFs as real data to train diffusion models. In contrast, we propose a new single-stage training paradigm with an end-to-end objective that jointly optimizes a NeRF auto-decoder and a latent diffusion model, enabling simultaneous 3D reconstruction and prior learning, even from sparsely available views. At test time, we can directly sample the diffusion prior for unconditional generation, or combine it with arbitrary observations of unseen objects for NeRF reconstruction. SSDNeRF demonstrates robust results comparable to or better than leading task-specific methods in unconditional generation and single/sparse-view 3D reconstruction.6 Hansheng Chen 0001, Jiatao Gu, Anpei Chen, Wei Tian 0001, Zhuowen Tu, Lingjie Liu, Hao Su 0001 |
ICCV | 4 |
| 2022 | EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose EstimationabstractLocating 3D objects from a single RGB image via Perspective-n-Points (PnP) is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest interpreting PnP as a differentiable layer, so that 2D-3D point correspondences can be partly learned by backpropagating the gradient w.r.t. object pose. Yet, learning the entire set of unrestricted 2D-3D points from scratch fails to converge with existing approaches, since the deterministic pose is inherently non-differentiable. In this paper, we propose the EPro-PnP a probabilistic PnP layer for general end-to-end pose estimation, which outputs a distribution of pose on the SE(3) manifold, essentially bringing categorical Softmax to the continuous domain. The 2D-3D coordinates and corresponding weights are treated as intermediate variables learned by minimizing the KL divergence between the predicted and target pose distribution. The underlying principle unifies the existing approaches and resembles the attention mechanism. EPro-PnP significantly outperforms competitive baselines, closing the gap between PnP-based method and the task-specific leaders on the LineMOD 6DoF pose estimation and nuScenes 3D object detection benchmarks.3 Hansheng Chen 0001, Pichao Wang, Fan Wang 0019, Wei Tian 0001, Lu Xiong 0001, Hao Li 0030 |
CVPR | 4 |
| 2021 | MonoRUn: Monocular 3D Object Detection by Reconstruction and Uncertainty PropagationabstractObject localization in 3D space is a challenging aspect in monocular 3D object detection. Recent advances in 6DoF pose estimation have shown that predicting dense 2D-3D correspondence maps between image and object 3D model and then estimating object pose via Perspective-n-Point (PnP) algorithm can achieve remarkable localization accuracy. Yet these methods rely on training with ground truth of object geometry, which is difficult to acquire in real outdoor scenes. To address this issue, we propose MonoRUn, a novel detection framework that learns dense correspondences and geometry in a self-supervised manner, with simple 3D bounding box annotations. To regress the pixel-related 3D object coordinates, we employ a regional reconstruction network with uncertainty awareness. For self-supervised training, the predicted 3D coordinates are projected back to the image plane. A Robust KL loss is proposed to minimize the uncertainty-weighted reprojection error. During testing phase, we exploit the network uncertainty by propagating it through all downstream modules. More specifically, the uncertainty-driven PnP algorithm is leveraged to estimate object pose and its covariance. Extensive experiments demonstrate that our proposed approach outperforms current state-of-the-art methods on KITTI benchmark.1 Hansheng Chen 0001, Yuyao Huang 0001, Wei Tian 0001, Zhong Gao, Lu Xiong 0001 |
CVPR | 3 |
| 2021 | Pedestrian Detection by Fusion of RGB and Infrared Images in Low-Light Environment
Qing Deng, Wei Tian 0001, Yuyao Huang 0001, Lu Xiong 0001 |
FUSION | 2 |
| 2020 | Online Multi-Object Tracking Using Joint Domain Information in Traffic ScenariosabstractVisual tracking of multiple objects is an essential component for a perception system in autonomous driving vehicles. One of the favorable approaches is the tracking-by-detection paradigm, which links current detection hypotheses to previously estimated object trajectories (also known as tracks) by searching appearance or motion similarities between them. As this search operation is usually based on a very limited spatial or temporal locality, the association can fail in cases of motion noise or long-term occlusion. In this paper, we propose a novel tracking method that solves this problem by putting together information from both enlarged structural and temporal domain. For efficiency without loss of optimality, this approach is decomposed in to three stages, with each dealing with only one constrained association task, and thus, it follows the alternating optimization fashion. In our approach, detections are first assembled into small tracklets based on meta-measurements of object affinity. The association task for tracklets-to-tracks is solved by structural information based on a motion pattern between them. Here, we propose new rules to decouple the processing time from the tracklet length. Furthermore, constraints from temporal domain are introduced to recover objects, which are long-time disappearing due to failed detection or long-term occlusion. By putting together the heterogeneous domain information, our approach exhibits an improved state-of-the-art performance on standard benchmarks. With relatively little processing time, an online and real-time tracking is also permitted in our approach. Wei Tian 0001, Martin Lauer, Long Chen 0005 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | A Collaborative Visual Tracking Architecture for Correlation Filter and Convolutional Neural Network LearningabstractVisual object tracking has achieved remarkable progress in recent years and has been broadly applied in intelligent transportation systems such as autonomous vehicles and drones to monitor and analyze the behavior of specific targets. One typical tracking approach is the discriminative tracker, which branches into two main categories: the correlation filter (CF) and the convolutional neural network (CNN). However, most of the current researches consider both categories as two separate techniques and only rely on one of them. Thus, a dense cooperation between the CF and the CNN still remains less discovered and the question of how to effectively join both techniques to further boost the tracking performance is still open. To address this issue, in this paper, we propose a collaborative architecture which incorporates models constructed with both techniques and dynamically aggregates their response maps for target inference. By an alternating optimization, both models are learned on each other's errors to persistently improve the classification power of the whole tracker. For further efficiency, we present a faster solver for our utilized CF and an analytical solution for dynamic model weighting. Through experiments on standard benchmarks, we reveal the influence of key factors on the joint learning architecture and show that it outperforms the state-of-the-art approaches. Wei Tian 0001, Niels Ole Salscheider, Yunxiao Shan, Long Chen 0005, Martin Lauer |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Toward the Ghosting Phenomenon in a Stereo-Based Map With a Collaborative RGB-D RepairabstractAlthough 3-D reconstruction of dynamic road environment by moving cameras has been broadly applied in recognition and navigation systems, this task is still considered challenging, especially under circumstances with moving objects, where the reconstruction precision is strongly harassed by the ghosting problem. To address this issue, in this paper, we propose a novel approach for reconstructing 3-D maps of complete static scenes, based on a combination of an elaborately designed moving-object filtering mechanism and a map repairing and blank refilling procedure, where both plausible color and depth information from stereo image pairs are utilized. In this approach, first, we employ the planarity knowledge into the initial depth map based on the simple linear iterative cluster (SLIC) superpixel segmentation. The dynamic area in the image is determined under the supervision of odometry calculation. After wiping off moving objects, by collaboratively repairing color and depth information, the final 3-D map containing only static scene is obtained. The experimental results on extensive challenging real-world scenarios demonstrate the effectiveness and robustness of our approach. Jiasong Zhu, Lei Fan 0005, Wei Tian 0001, Long Chen 0005, Dongpu Cao, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | High-Resolution Driving Scene Synthesis Using Stacked Conditional Gans and Spectral NormalizationabstractLarge-scale dataset plays a key role in the driving scene understanding for deep learning based-autonomous driving tasks. Due to the fact that the annotation for a large number of images is extremely labor-intensive and time-consuming, many researchers turn to using image-synthesis techniques for automatic construction of training data. However, traditional methods often have difficulties in producing high-definition driving scene images. To tackle this problem, in this paper, we propose a novel deep model - hdCGAN - for high-definition image-to-image translation. The hdCGAN is built on a conditional GAN in combination with a spectral normalization. Moreover, we improve the hdCGAN by using a stacked network architecture and the enhanced model is called stack-hdCGAN. With the guidance of multi-scale discriminators and the constraint of spectral normalization in the training procedure, the learned models can generate high-resolution and high-quality driving scene images from corresponding semantic segmentation maps. Quantitative and qualitative evaluations on the Cityscapes dataset demonstrate the effectiveness of the proposed models. Shaobo Lin, Long Chen 0005, Qin Zou 0001, Wei Tian 0001 |
ICME | 4 |
| 2019 | Improving classification with semi-supervised and fine-grained learning
Danyu Lai, Wei Tian 0001, Long Chen 0005 |
Pattern Recognit. | 2 |
| 2019 | Deep Integration: A Multi-Label Architecture for Road Scene RecognitionabstractDeep convolutional neural networks have been applied by automobile industries, Internet giants, and academic institutes to boost autonomous driving technologies; while progress has been witnessed in environmental perception tasks, such as object detection and driver state recognition, the scene-centric understanding and identification still remain a virgin land. This mainly encompasses two key issues: 1) the lack of shared large datasets with comprehensively annotated road scene information and 2) the difficulty to find effective ways to train networks concerning the bias of category samples, image resolutions, scene dynamics, and capturing conditions. In this paper, we make two contributions: 1) we introduce a large-scale dataset with over 110 k images, dubbed DrivingScene, covering traffic scenarios under different weather conditions, road structures, and environmental instances and driving places, which is the first large-scale dataset for multi-class traffic scenes classification and 2) we propose a multi-label neural network for road scene recognition, which incorporates both single- and multi-class classification modes into a multi-level cost function for training with imbalanced categories and utilizes a deep data integration strategy to improve the classification ability on hard samples. The experimental results on DrivingScene and PASCAL VOC demonstrate the effectiveness of the proposed approach in handling the challenge of data imbalance. Long Chen 0005, Wujing Zhan, Wei Tian 0001, Qin Zou 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Vehicle Tracking at Nighttime by Kernelized Experts With Channel-Wise and Temporal Reliability EstimationabstractDespite the fact that in recent years, vision-based tracking approaches have made significant progress, the task of tracking vehicles at night still remains challenging. Visual information is strongly deteriorated or at least degraded due to poor illumination conditions. This reduces the perceptive ability of vision systems significantly and can even lead to target loss, resulting in false estimation and/or false prediction of object behavior. In this paper, we propose a novel online-learning method to track vehicles at night. Our method is based on the kernelized correlation filter and assembles different feature channels to kernelized experts. By estimating their reliabilities, we force the appearance model to focus on the most discriminative visual features to accomplish the classification. In addition, a temporal optimization step in conjunction with a memory model is used to remove outliers and keep the most reliable samples to train the tracker models. Experiments over various daytime and weather conditions show that our approach outperforms existing trackers at night and in case of bad weather while offering state-of-the-art performance in more favorable situations. As our tracker has only little computational cost, it is appropriate for use cases with real-time requirements like in automotive or industrial applications. Wei Tian 0001, Long Chen 0005, Ke Zou, Martin Lauer |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | UA-DETRAC 2017: Report of AVSS2017 & IWT4S Challenge on Advanced Traffic MonitoringabstractThe rapid advances of transportation infrastructure have led to a dramatic increase in the demand for smart systems capable of monitoring traffic and street safety. Fundamental to these applications are a community-based evaluation platform and benchmark for object detection and multi-object tracking. To this end, we organize the AVSS2017 Challenge on Advanced Traffic Monitoring, in conjunction with the International Workshop on Traffic and Street Surveillance for Safety and Security (IWT4S), to evaluate the state-of-the-art object detection and multi-object tracking algorithms in the relevance of traffic surveillance. Submitted algorithms are evaluated using the large-scale UA-DETRAC benchmark and evaluation protocol. The benchmark, the evaluation toolkit and the algorithm performance are publicly available from the website http://detrac-db.rit.albany.edu. Siwei Lyu, Ming-Ching Chang, Dawei Du, Longyin Wen, Honggang Qi, Yuezun Li, Yi Wei 0006, Lipeng Ke, Tao Hu 0011, Marco Del Coco, Pierluigi Carcagnì, Dmitriy Anisimov, Erik Bochinski, Fabio Galasso, Filiz Bunyak, Hao Ye 0005, Hong Wang 0014, Kannappan Palaniappan, Koray Ozcan, Li Wang 0033, Liang Wang 0001, Martin Lauer, Nattachai Watcharapinchai, Nenghui Song, Noor Al-Shakarji, Sikandar Amin, Sitapa Watcharapinchai, Tatiana Khanova, Thomas Sikora, Tino Kutschbach, Volker Eiselein, Wei Tian 0001, Xiangyang Xue 0001, Xiaoyi Yu, Yao Lu 0028, Yingbin Zheng, Yongzhen Huang, Yuqi Zhang 0001 |
AVSS | 34 |
| 2017 | Joint tracking with event grouping and temporal constraintsabstractVision systems become more and more popular to be applied in monitoring tasks such as controlling traffic flows or for security issues. The analysis of target behavior is always based on its observed trajectory, which can be acquired by tracking approaches. Although the fashion of tracking-by-detection is favored by the research community, it still faces challenges like unexpected occlusion caused by background objects or other tracked targets, which can interfere the matching operation and result in tracking errors. In this paper, we propose a novel approach by aggregating prediction events within target groups and integrating a graph-modeling based stitching procedure to handle the above mentioned problems. The evaluation results on the UA-DETRAC benchmark demonstrated the state-of-the-art performance of our tracking approach. Wei Tian 0001, Martin Lauer |
AVSS | 1 |