VLDB 2026 Research / reviewers in the wild / expert
Yang Chen 0037
dblp:48/4792-37
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0001-8187-2000ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MECI: A Multi-Model Motion Capture-Free Event Dataset Featuring Large-Scale Challenging Indoor EnvironmentsabstractRecently, several event-based datasets have emerged to foster the application of the new event camera to classic vision tasks like Simultaneous Localization and Mapping (SLAM). However, current indoor benchmark datasets depend on the expensive motion capture system to obtain ground-truth trajectories, restricting data acquisition to small object-centric scenes or single-room environments due to infrastructure costs and spatial limitations. Furthermore, these datasets lack sensor diversity, relying solely on a single event camera model that hinders practical cross-device generalization. To address the above limitations, we propose MECI, the first Multi-model Event dataset targeting Challenging Indoor environments, especially including large-scale scenes with complete trajectory ground-truth provided. Specifically, MECI includes 38 posed sequences of visual-inertial-event data from three different models of event cameras with varying resolutions and data frequencies. Apart from typical challenging factors such as illumination changes, motion blur, and dynamics, these sequences involve large-scale indoor scenes (across rooms and floors) with each room occupying approximately 60 m \({}^{2}\) and a maximum trajectory length of 343.8 m. We innovatively leverage ETS (Electronic Total Station) measurements and AprilTag markers to provide complete 6-DoF trajectories with low cost and minimal environmental modifications. Extensive experiments show that our MECI is an effective yet challenging benchmark dataset not only for visual SLAM, but also for event-image reconstruction. Our project page is https://cslinzhang.github.io/MECI_dataset/ . Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | WSGS: A Speech-Driven Zero-Shot System for 6D Robotic Arm GraspingabstractIn the robotic vision industry, recent years have witnessed a growing interest in object detection and pose estimation for precise robotic arm grasping. In fact, various unfavorable factors, such as the size limitations of robotic grippers, the diversity and potential complexity of object shapes and poses, and the cluttered environment, make robotic arm grasping based on 6D object poses much harder than it seems. In this paper, to solve these issues to some extent, we proposed a speech-driven zero-shot system for robotic arm grasping, called WSGS (Whisper-SAM6D Grasping System). It enables speech-driven human-interactive grasping with the Franka Emika robotic arm by performing instance segmentation and pose estimation based on speech instructions. Specifically, WSGS accurately recognizes the 6D pose of an unknown object and adapts to its size to find the most suitable position for grasping. Comprehensive experiments on our real-world scenarios demonstrate that WSGS can produce high-accuracy instance segmentation and pose estimation results, achieving adaptive robotic arm grasping of unknown objects based on speech commands. Yitong Ge, Lin Zhang 0014, Yang Chen 0037, Ying Shen 0005 |
ICME | 3 |
| 2025 | GRE-SLAM: 6-DoF Pure Event-Based SLAM with Semi-Dense Depth Recovery Assisted Bundle AdjustmentabstractEvent cameras are innovative bioinspired vision sensors that output pixel-level brightness changes instead of standard intensity frames. Such cameras do not suffer from motion blur and cope well with scenes characterized by high dynamic range, which can benefit classic computer vision tasks such as pose estimation. However, currently developed event-based pose estimation methods either require extra data as inputs (such as IMU data or depths) or lack a global refinement step to alleviate accumulated drifts. To this end, we propose the first 6-DoF pure event-based SLAM system equipped with back-end global optimization, named GRE-SLAM (Globally Refined Event-based SLAM). For robustness and accuracy, first, 6-DoF motion compensation is introduced in the front-end to prepare sharp-edged event frames and a favorable initialization pose, mitigating unstable optimization during event registration brought by sparsity and noise of events. Second, a novel adaptive semi-dense depth recovery algorithm enriches front-end's sparse depths without additional sensors, helping establish long-term edge alignment constraints to support global BA in the back-end. Comprehensive experiments on real-world datasets demonstrate that our method can produce high-accuracy pose estimation results as well as recover a semi-dense depth map for each Image of Warped Events (IWE). Yang Chen 0037, Lin Zhang 0014 |
ICMR | 1 |
| 2025 | Online indoor visual odometry with semantic assistance under implicit epipolar constraints
Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
Pattern Recognit. | 1 |
| 2025 | ATM-NeRF: Accelerating Training for NeRF Rendering on Mobile Devices via Geometric RegularizationabstractRecently, an increasing number of researchers have been dedicated to transferring the impressive novel view synthesis capability of Neural Radiance Fields (NeRF) to resource-constrained mobile devices. One common solution is to pre-train NeRF and bake it into textured meshes which are well supported by mobile graphics hardware. However, the training process of existing methods often requires several hours even with multiple high-end NVIDIA V100 GPUs. The underlying reason is that these schemes mainly rely on photometric rendering loss, neglecting the geometric relationship between the pre-trained NeRF and the baked results. Standing on this point, we presentATM-NeRF(AcceleratingTraining forMobile rendering based onNeRF), which is the first to apply effective geometric regularization constraints during both the pre-training and the baking training stages for faster convergence. Specifically, in the initial NeRF pre-training stage, we enforce consistency of the multi-resolution density grids representing the scene geometry to mitigate the shape-radiance ambiguity problem to some extent, achieving a coarse mesh with smoothness. In the second stage, we utilize the positions and geometric features of 3D points projected from the pre-trained posed depths to provide geometric supervision for joint refinement of geometry and appearance of the coarse mesh. As a result, our ATM-NeRF achieves comparable rendering quality to MobileNeRF with a training speed that is about$30\times \sim 70\times$faster while maintaining finer structure details of the exported mesh. Yang Chen 0037, Lin Zhang 0014, Shengjie Zhao 0001, Yicong Zhou |
IEEE Trans. Multim. | 1 |
| 2024 | I2P Registration by Learning the Underlying Alignment Feature Space from Pixel-to-Point SimilaritiesabstractEstimating the relative pose between a camera and a LiDAR holds paramount importance in facilitating complex task execution within multi-agent systems. Nonetheless, current methodologies encounter two primary limitations. First, amid the cross-modal feature extraction, they typically employ separate modal branches to extract cross-modal features from images and point clouds. This approach results in the feature spaces of images and point clouds being misaligned, thereby reducing the robustness of establishing correspondences. Second, due to the scale differences between images and point clouds, one-to-many pixel-point correspondences are inevitably encountered, which will mislead the pose optimization. To address these challenges, we propose a framework named I mage-to- P oint cloud registration by learning the underlying alignment feature space from P ixel-to- P oint SIM imilarities (I2P \({}_{\mathbf{ppsim}}\) ) . Central to \(\text{I2P}_{\text{ppsim}}\) is a Shared Feature Alignment Module (SFAM). It is designed under on a coarse-to-fine architecture and uses a weight-sharing network to construct an alignment feature space. Benefiting from SFAM, \(\text{I2P}_{\text{ppsim}}\) can effectively identify the co-view regions between images and point clouds and establish high-reliability 2D-3D correspondences. Moreover, to mitigate the one-to-many correspondence issue, we introduce a similarity maximization strategy termed point-max. This strategy effectively filters out outliers, thereby establishing accurate 2D-3D correspondences. To evaluate the efficacy of our framework, we conduct extensive experiments on KITTI Odometry and Oxford Robotcar. The results corroborate the effectiveness of our framework in improving image-to-point cloud registration. To make our results reproducible, the source codes have been released at https://cslinzhang.github.io/I2P Yunda Sun, Lin Zhang 0014, Zhong Wang 0009, Yang Chen 0037, Shengjie Zhao 0001, Yicong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Extrinsic Self-Calibration of the Surround-View System: A Weakly Supervised ApproachabstractAn SVS usually consists of four wide-angle fisheye cameras mounted around the vehicle to sense the surrounding environment. From the images synchronously captured by all cameras, a top-down surround-view can be synthesized, on the premise that both intrinsics and extrinsics of the cameras have been calibrated. At present, the intrinsic calibration approach is relatively well-developed and can be pipelined, while the extrinsic calibration is still immature. On one hand, the existing manual calibration schemes are usually reliable, but need to be conducted by professionals in specific sites, which is undoubtedly cumbersome. On the other hand, the majority of the existing self- calibration schemes are based on low-level features and their stability and robustness are usually unsatisfactory. As far as we know, an effective extrinsic self-calibration scheme designed specially for the SVS is still lacking. To fill such a research gap to some extent, we propose a novel self-calibration scheme which follows a weakly supervised framework, namely WESNet (Weakly-supervised Extrinsic Self-calibration Network). The training of WESNet consists of two stages. First, we utilize the corners in a few calibration site images as the weak supervision to roughly optimize the network by minimizing the geometric loss. Then, after the convergency in the first stage, we additionally introduce a self-supervised photometric loss term that can be constructed by the photometric information from natural images for further fine-tuning. Besides, to support training, we totally collected 19,078 groups of synchronously captured fisheye images under various environmental conditions. To our knowledge, thus far this is the largest surround-view dataset containing original fisheye images. By means of learning prior knowledge from the training data, WESNet takes the original fisheye images synchronously collected as the input, and directly yields extrinsics end-to-end with little labor cost. Its efficiency and efficacy have been corroborated by extensive experiments conducted on our collected dataset. To make our results reproducible, source code and the collected dataset have been released at https://cslinzhang.github.io/WESNet/WESNet.html. Yang Chen 0037, Lin Zhang 0014, Ying Shen 0005, Brian Nlong Zhao, Yicong Zhou |
IEEE Trans. Multim. | 1 |
| 2022 | CVIDS: A Collaborative Localization and Dense Mapping Framework for Multi-Agent Based Visual-Inertial SLAMabstractNowadays, visual SLAM (Simultaneous Localization And Mapping) has become a hot research topic due to its low costs and wide application scopes. Traditional visual SLAM frameworks are usually designed for single-agent systems, completing both the localization and the mapping with sensors equipped on a single robot or a mobile device. However, the mobility and work capacity of the single agent are usually limited. In reality, robots or mobile devices sometimes may be deployed in the form of clusters, such as drone formations, wearable motion capture systems, and so on. As far as we know, existing SLAM systems designed for multi-agents are still sporadic, and most of them have non-negligible limitations in functions. Specifically, on one hand, most of the existing multi-agent SLAM systems can only extract some key features and build sparse maps. On the other hand, schemes that can reconstruct the environment densely cannot get rid of the dependence on depth sensors, such as RGBD cameras or LiDARs. Systems that can yield high-density maps just with monocular camera suites are temporarily lacking. As an attempt to fill in the research gap to some extent, we design a novel collaborative SLAM system, namely CVIDS (Collaborative Visual-Inertial Dense SLAM), which follows a centralized and loosely coupled framework and can be integrated with any existing Visual-Inertial Odometry (VIO) to accomplish the co-localization and the dense reconstruction. Integrating our proposed robust loop closure detection module and two-stage pose-graph optimization pipeline, the co-localization module of CVIDS can estimate the poses of different agents in a unified coordinate system efficiently from the packed images and local poses sent by the client-ends of different agents. Besides, our motion-based dense mapping module can effectively recover the 3D structures of selected keyframes and then fuse their depth information to the global map for reconstruction. The superior performance of CVIDS is corroborated by both quantitative and qualitative experimental results. To make our results reproducible, the source code has been released at https://cslinzhang.github.io/CVIDS. Tianjun Zhang, Lin Zhang 0014, Yang Chen 0037, Yicong Zhou |
IEEE Trans. Image Process. | 3 |