EDBT 2026 Demo / reviewers in the wild / expert
Wei Chen 0092
dblp:181/2832-92
· DBLP profile ↗
14ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-6314-5600ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MHED-SLAM: Multi-Scale Hybrid Encoding-Based Decoupled SLAMabstractNeural Radiance Fields (NeRF)-based Visual Simultaneous Localization and Mapping (SLAM) achieve superior scene geometric modeling and robust camera tracking by leveraging neural representations. Existing methods typically relied on multi-resolution hash encoding with truncated signed distance fields (TSDF) to achieve high frame rates. However, unavoidable hash collisions can lead to artifacts, and multi-view color inconsistencies in indoor scenes can result in shape-radiance ambiguity, adversely affecting geometric quality and tracking accuracy. To address these issues, we propose a novel Multi-scale Hybrid Encoding-based Decoupled SLAM (MHED-SLAM). First, to mitigate the adverse effects of hash collisions and reduce the number of learnable parameters, we innovatively fuse a coarse-scale hash tri-plane with a fine-scale hash grid within a single latent volume. Second, to enable precise geometric reconstruction and camera tracking, we decouple the reconstruction and rendering processes, independently learning a TSDF field for reconstruction and a density field for rendering. Third, we devise a Symmetric Kullback-Leibler (SKL) strategy based on ray termination distributions to align the probability distributions derived from the TSDF and density fields for their synchronous convergence. Extensive experimental evaluations demonstrate that our approach surpasses the state-of-the-art (SOTA) methods by utilizing a faster frame rate of 20 Hz and fewer parameters, while achieving higher tracking and reconstruction accuracy. Dengfang Feng, Wenyang Qin, Zhongchen Shi, Wei Chen 0092, Yanhui Duan, Liang Xie 0012, Erwei Yin |
AAAI | 4 |
| 2026 | Locomotion in CAVE: Enhancing immersion through full-body motion
Zhongchen Shi, Wei Chen 0092, Liang Xie 0012, Meng Gai, Suxia Zhang, Erwei Yin |
Comput. Graph. | 4 |
| 2025 | M2EIT: Multi-Domain Mixture of Experts for Robust Neural Inertial Tracking
Changhao Chen, Zhongchen Shi, Wei Chen 0092, Liang Xie 0012, Erwei Yin |
ICCV | 5 |
| 2024 | Trajectory-based Calibration for Optical See-Through Head-Mounted Displays Without Alignment
Shaohua Zhao, Wei Chen 0092, Zhongchen Shi, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
PRCV (6) | 3 |
| 2024 | MVINS: Tightly Coupled Mocap-Visual-Inertial Fusion for Global and Drift-Free Pose EstimationabstractAugmented reality (AR), a prominent application within the Internet of Things (IoT) domain, demands high-performance pose estimation. Presently, the visual-inertial navigation system (VINS) is acknowledged as an essential method for providing 6-DoF poses. However, VINS builds the local frame at random during the system initialization stage, making it difficult to establish a connection with the global frame. In addition, VINS is prone to drifting. In this paper, we propose an innovative method that tightly couples markerless motion capture (Mocap) with vision and an IMU to achieve global and drift-free pose estimation for AR glasses. To address the issue of pose initialization and establish a connection between the IMU and Mocap, we introduce a coarse-to-fine initialization strategy, enabling data fusion for Mocap, vision, and the IMU under a unified global frame. Furthermore, we formulate the Mocap factor alongside the visual and inertial factors and integrate them into a factor graph framework to constrain the system states. With a spatiotemporal calibration method, the IMU-Mocap extrinsic parameter and time offset are calibrated online to improve the pose estimation accuracy. Experimental evaluations in real-world experiments demonstrate the capability of our method to accurately estimate drift-free poses in the global frame. Compared to the state-of-the-art VINS-Fusion, ORB-SLAM3, and GVIS, we achieve improvements of 81%, 42%, and 33% in translation accuracy and improvements of 58%, 33%, and 72% in rotation accuracy, respectively. Moreover, we also evaluate our system for the EuRoC dataset, further indicating the effectiveness of the proposed work. Liang Xie 0012, Wei Wang 0076, Zhongchen Shi, Wei Chen 0092, Ye Yan 0001, Erwei Yin |
IEEE Internet Things J. | 5 |
| 2024 | Trajectory-based alignment for optical see-through HMD calibrationabstractAbstract In order to align the virtual and real content precisely through augmented reality devices, especially in optical see-through head-mounted displays (OST-HMD), it is necessary to calibrate the device before using it. However, most existing methods estimated the parameters via 3D-2D correspondences based on the 2D alignment, which is cumbersome, time-consuming, theoretically complex, and results in insufficient robustness. To alleviate this issue, in this paper, we propose an efficient and simple calibration method based on the principle of directly calculating the projection transformation between virtual space and the real world via 3D-3D alignment. The proposed method merely needs to record the motion trajectory of the cube-marker in the real and virtual world, and then calculate the transformation matrix between the virtual space and the real world by aligning the two trajectories in the observed view. There are two advantages associated with the proposed method. First, the operation is simple. Theoretically, the user only needs to perform four alignment operations for calibration without changing the rotation variation. Second, the trajectory can be easily distributed throughout the entire observation view, resulting in more robust calibration results. To validate the effectiveness of the proposed method, we conducted extensive experiments on our self-built optical see-through head-mounted display (OST-HMD) device. The experimental results show that the proposed method can achieve better calibration results than other calibration methods. Lingling Chen, Shaohua Zhao, Wei Chen 0092, Zhongchen Shi, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
Multim. Tools Appl. | 3 |
| 2023 | Fourier-Net: Fast Image Registration with Band-Limited DeformationabstractUnsupervised image registration commonly adopts U-Net style networks to predict dense displacement fields in the full-resolution spatial domain. For high-resolution volumetric image data, this process is however resource-intensive and time-consuming. To tackle this problem, we propose the Fourier-Net, replacing the expansive path in a U-Net style network with a parameter-free model-driven decoder. Specifically, instead of our Fourier-Net learning to output a full-resolution displacement field in the spatial domain, we learn its low-dimensional representation in a band-limited Fourier domain. This representation is then decoded by our devised model-driven decoder (consisting of a zero padding layer and an inverse discrete Fourier transform layer) to the dense, full-resolution displacement field in the spatial domain. These changes allow our unsupervised Fourier-Net to contain fewer parameters and computational operations, resulting in faster inference speeds. Fourier-Net is then evaluated on two public 3D brain datasets against various state-of-the-art approaches. For example, when compared to a recent transformer-based method, named TransMorph, our Fourier-Net, which only uses 2.2% of its parameters and 6.66% of the multiply-add operations, achieves a 0.5% higher Dice score and an 11.48 times faster inference speed. Code is available at https://github.com/xi-jia/Fourier-Net. Xi Jia, Joseph Bartlett, Wei Chen 0092, Siyang Song, Tianyang Miller, Xinxing Cheng, Wenqi Lu 0001, Zhaowen Qiu, Jinming Duan 0001 |
AAAI | 3 |
| 2023 | One-Stage Wireframe Parsing in Fish-Eye Images
Ruqiang Huang, Zhongchen Shi, Wei Chen 0092, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
PRCV (11) | 4 |
| 2022 | Learning a Model-Driven Variational Network for Deformable Image RegistrationabstractData-driven deep learning approaches to image registration can be less accurate than conventional iterative approaches, especially when training data is limited. To address this issue and meanwhile retain the fast inference speed of deep learning, we propose VR-Net, a novel cascaded variational network for unsupervised deformable image registration. Using a variable splitting optimization scheme, we first convert the image registration problem, established in a generic variational framework, into two sub-problems, one with a point-wise, closed-form solution and the other one being a denoising problem. We then propose two neural layers (i.e. warping layer and intensity consistency layer) to model the analytical solution and a residual U-Net (termed generalized denoising layer) to formulate the denoising problem. Finally, we cascade the three neural layers multiple times to form our VR-Net. Extensive experiments on three (two 2D and one 3D) cardiac magnetic resonance imaging datasets show that VR-Net outperforms state-of-the-art deep learning methods on registration accuracy, whilst maintaining the fast inference speed of deep learning and the data-efficiency of variational models. Xi Jia, Alexander Thorley, Wei Chen 0092, Huaqi Qiu, LinLin Shen, Iain B. Styles, Hyung Jin Chang, Ales Leonardis, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert, Jinming Duan 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2021 | FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation MechanismabstractIn this paper, we focus on category-level 6D pose and size estimation from a monocular RGB-D image. Previous methods suffer from inefficient category-level pose feature extraction, which leads to low accuracy and inference speed. To tackle this problem, we propose a fast shape-based network (FS-Net) with efficient category-level feature extraction for 6D pose estimation. First, we design an orientation aware autoencoder with 3D graph convolution for latent feature extraction. Thanks to the shift and scale-invariance properties of 3D graph convolution, the learned latent feature is insensitive to point shift and object size. Then, to efficiently decode category-level rotation information from the latent feature, we propose a novel decoupled rotation mechanism that employs two decoders to complementarily access the rotation information. For translation and size, we estimate them by two residuals: the difference between the mean of object points and ground truth translation, and the difference between the mean size of the category and ground truth size, respectively. Finally, to increase the generalization ability of the FS-Net, we propose an on-line box-cage based 3D deformation mechanism to augment the training data. Extensive experiments on two benchmark datasets show that the proposed method achieves state-of-the-art performance in both category- and instance-level 6D object pose estimation. Especially in category-level pose estimation, without extra synthetic data, our method outperforms existing methods by 6.3% on the NOCS-REAL dataset1. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, LinLin Shen, Ales Leonardis |
CVPR | 1 |
| 2021 | PointFace: Point Set Based Feature Learning for 3D Face RecognitionabstractThough 2D face recognition (FR) has achieved great success due to powerful 2D CNNs and large-scale training data, it is still challenged by extreme poses and illumination conditions. On the other hand, 3D FR has the potential to deal with aforementioned challenges in the 2D domain. However, most of available 3D FR works transform 3D surfaces to 2D maps and utilize 2D CNNs to extract features. The works directly processing point clouds for 3D FR is very limited in literature. To bridge this gap, in this paper, we propose a light-weight framework, named PointFace, to directly process point set data for 3D FR. Inspired by contrastive learning, our PointFace use two weight-shared encoders to directly extract features from a pair of 3D faces. A feature similarity loss is designed to guide the encoders to obtain discriminative face representations. We also present a pair selection strategy to generate positive and negative pairs to boost training. Extensive experiments on Lock3DFace and Bosphorus show that the proposed PointFace outperforms state-of-the-art 2D CNN based methods. Changyuan Jiang, Shisong Lin, Wei Chen 0092, Feng Liu 0013, LinLin Shen |
IJCB | 3 |
| 2020 | G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector FeaturesabstractIn this paper, we propose a novel real-time 6D object pose estimation framework, named G2L-Net. Our network operates on point clouds from RGB-D detection in a divide-and-conquer fashion. Specifically, our network consists of three steps. First, we extract the coarse object point cloud from the RGB-D image by 2D detection. Second, we feed the coarse object point cloud to a translation localization network to perform 3D segmentation and object translation prediction. Third, via the predicted segmentation and translation, we transfer the fine object point cloud into a local canonical coordinate, in which we train a rotation localization network to estimate initial object rotation. In the third step, we define point-wise embedding vector features to capture viewpoint-aware information. To calculate more accurate rotation, we adopt a rotation residual estimator to estimate the residual between initial rotation and ground truth, which can boost initial pose estimation performance. Our proposed G2L-Net is real-time despite the fact multiple steps are stacked via the proposed coarse-to-fine framework. Extensive experiments on two benchmark datasets show that G2L-Net achieves state-of-the-art performance in terms of both accuracy and speed. Wei Chen 0092, Xi Jia, Hyung Jin Chang, Jinming Duan 0001, Ales Leonardis |
CVPR | 1 |
| 2020 | One-stage Multi-task Detector for 3D Cardiac MR ImagingabstractFast and accurate landmark location and bounding box detection are important steps in 3D medical imaging. In this paper, we propose a novel multi-task learning framework, for real-time, simultaneous landmark location and bounding box detection in 3D space. Our method extends the famous single-shot multibox detector (SSD) from single-task learning to multitask learning and from 2D to 3D. Furthermore, we propose a post-processing approach to refine the network landmark output, by averaging the candidate landmarks. Owing to these settings, the proposed framework is fast and accurate. For 3D cardiac magnetic resonance (MR) images with size 224×224×64, our framework runs ~128 volumes per second (VPS) on GPU and achieves 6.75mm average point-to-point distance error for landmark location, which outperforms both state-of-the-art and baseline methods. We also show that segmenting the 3D image cropped with the bounding box results in both improved performance and efficiency. Weizeng Lu, Xi Jia, Wei Chen 0092, Nicolò Savioli, Antonio M. Simoes Monteiro de Marvao, LinLin Shen, Declan P. O'Regan, Jinming Duan 0001 |
ICPR | 3 |
| 2020 | PointPoseNet: Point Pose Network for Robust 6D Object Pose EstimationabstractIn this paper, we propose a novel pipeline to estimate 6D object pose from RGB-D images of known objects present in complex scenes. The pipeline directly operates on raw point clouds extracted from RGB-D scans. Specifically, our method takes the point cloud as input and regresses the point-wise unit vectors pointing to the 3D keypoints. We then use these vectors to generate keypoint hypotheses from which the 6D object pose hypotheses are computed. Finally, we select the best 6D object pose from the hypotheses based on a proposed scoring mechanism with geometry constraints. Extensive experiments show that the proposed method is robust against the variety in object shape and appearance as well as occlusions between objects, and that our method outperforms the state-of-the-art methods on the LINEMOD and Occlusion LINEMOD datasets. Wei Chen 0092, Jinming Duan 0001, Hector Basevi, Hyung Jin Chang, Ales Leonardis |
WACV | 1 |