EDBT 2026 Demo / reviewers in the wild / expert
Yi Zhou 0010
dblp:01/1901-10
· DBLP profile ↗
18ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0003-3201-8873ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Deep Representation Learning for Event-Enhanced Visual Autonomous Perception: The eAP DatasetabstractRecent visual autonomous perception systems achieve remarkable performances with deep representation learning. However, they fail in scenarios with challenging illumination. While event cameras can mitigate this problem, there is a lack of a large-scale dataset to develop event-enhanced deep visual perception models in autonomous driving scenes. To address the gap, we present theeAP(event-enhancedAutonomousPerception) dataset, the largest dataset with event cameras for autonomous perception. We demonstrate howeAPcan facilitate the study of different autonomous perception tasks, including 3D vehicle detection and object time-to-contact (TTC) estimation, through deep representation learning. Based oneAP, we demonstrate the first successful use of events to improve a popular 3D vehicle detection network in challenging illumination scenarios.eAPalso enables a devoted study of the representation learning problem of object TTC estimation. We show how a geometry-aware representation learning framework leads to the best event-based object TTC estimation network that operates at 200 FPS. The dataset, code, and pre-trained models will be made publicly available for future research. Shichao Li 0002, Qing Lian, Peiliang Li 0001, Xiaozhi Chen, Yi Zhou 0010 |
IEEE Trans. Robotics | 6 |
| 2025 | Convex Relaxation for Robust Vanishing Point Estimation in Manhattan WorldabstractDetermining the vanishing points (VPs) in a Manhattan world, as a fundamental task in many 3D vision applications, consists of jointly inferring the line-VP association and locating each VP. Existing methods are, however, either sub-optimal solvers or pursuing global optimality at a significant cost of computing time. In contrast to prior works, we introduce convex relaxation techniques to solve this task for the first time. Specifically, we employ a "soft" association scheme, realized via a truncated multi-selection error, that allows for joint estimation of VPs’ locations and line-VP associations. This approach leads to a primal problem that can be reformulated into a quadratically constrained quadratic programming (QCQP) problem, which is then relaxed into a convex semidefinite programming (SDP) problem. To solve this SDP problem efficiently, we present a globally optimal outlier-robust iterative solver (called GlobustVP), which independently searches for one VP and its associated lines in each iteration, treating other lines as outliers. After each independent update of all VPs, the mutual orthogonality between the three VPs in a Manhattan world is reinforced via local refinement. Extensive experiments on both synthetic and real-world data demonstrate that GlobustVP achieves a favorable balance between efficiency, robustness, and global optimality compared to previous works. The code is publicly available at github.com/wu-cvgl/GlobustVP. Bangyan Liao, Zhenjun Zhao, Haoang Li, Yi Zhou 0010, Yingping Zeng, Peidong Liu 0001 |
CVPR | 4 |
| 2025 | MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAMabstractRecent advancements in 3D Gaussian Splatting (3DGS) have made a significant impact on rendering and reconstruction techniques. Current research predominantly focuses on improving rendering performance and reconstruction quality using high-performance desktop GPUs, largely overlooking applications for embedded platforms like micro air vehicles (MAVs). These devices, with their limited computational resources and memory, often face a trade-off between system performance and reconstruction quality. In this paper, we improve existing methods in terms of GPU memory usage while enhancing rendering quality. Specifically, to address redundant 3D Gaussian primitives in SLAM, we propose merging them in voxel space based on geometric similarity. This reduces GPU memory usage without impacting system runtime performance. Furthermore, rendering quality is improved by initializing 3D Gaussian primitives via Patch-Grid (PG) point sampling, enabling more accurate modeling of the entire scene. Quantitative and qualitative evaluations on publicly available datasets demonstrate the effectiveness of our improvements. Yinlong Bai, Junkai Niu, Yijia He, Yi Zhou 0010 |
IROS | 7 |
| 2025 | E-MoFlow: Learning Egomotion and Optical Flow from Event Data via Implicit RegularizationabstractThe estimation of optical flow and 6-DoF ego-motion—two fundamental tasks in 3-D vision—has typically been addressed independently.
For neuromorphic vision (e.g., event cameras), however, the lack of robust data association makes solving the two problems separately an ill-posed challenge, especially in the absence of supervision via ground truth.
Existing works mitigate this ill-posedness by either enforcing the smoothness of the flow field via an explicit variational regularizer or leveraging explicit structure-and-motion priors in the parametrization to improve event alignment.
The former notably introduces bias in results and computational overhead, while the latter—which parametrizes the optical flow in terms of the scene depth and the camera motion—often converges to suboptimal local minima.
To address these issues, we propose an unsupervised pipeline that jointly optimizes egomotion and flow via implicit spatial-temporal and geometric regularization. First, by modeling camera's egomotion as a continuous spline and optical flow as an implicit neural representation, our method inherently embeds spatial-temporal coherence through inductive biases. Second, we incorporate structure-and-motion priors through differential geometric constraints, bypassing explicit depth estimation while maintaining rigorous geometric consistency.
As a result, our framework (called \textbf{E-MoFlow}) unifies egomotion and optical flow estimation via implicit regularization under a fully unsupervised paradigm. Experiments demonstrate its versatility to general 6-DoF motion scenarios, achieving state-of-the-art performance among unsupervised methods and competitive even with supervised approaches.
Code will be released upon acceptance. Wenpu Li, Bangyan Liao, Yi Zhou 0010, Pian Wan, Peidong Liu 0001 |
NeurIPS | 3 |
| 2025 | Event-Based Visual-Inertial State Estimation for High-Speed ManeuversabstractNeuromorphic event-based cameras are bio-inspired visual sensors with asynchronous pixels and extremely high temporal resolution. Such favorable properties make them an excellent choice for solving state estimation tasks under high-speed maneuvers. However, failures of camera pose tracking are frequently witnessed in state-of-the-art event-based visual odometry systems when the local map cannot be updated timely or feature matching is unreliable. One of the biggest roadblocks in this field is the absence of efficient and robust methods for data association without imposing any assumptions on the environment. This problem seems, however, unlikely to be addressed as in standard vision because of the motion-dependent nature of event data. To address this, we propose a map-free design for event-based visual-inertial state estimation in this paper. Instead of estimating camera position, we find that recovering the instantaneous linear velocity aligns better with event cameras' differential working principle. The proposed system uses raw data from a stereo event camera and an inertial measurement unit (IMU) as input, and adopts a dual-end architecture. The front-end preprocesses raw events and executes the computation of normal flow and depth information. To handle the temporally non-equispaced event data and establish association with temporally non-aligned IMU's measurements, the back-end employs a continuous-time formulation and a sliding-window scheme that can progressively estimate the linear velocity and IMU's bias. Experiments on synthetic and real data show our method achieves low-latency, metric-scale velocity estimation. To the best of our knowledge, this is the first real-time, purely event-based visual-inertial state estimator for high-speed maneuvers, requiring only sufficient textures and imposing no additional constraints on either the environment or motion pattern. Xiuyuan Lu, Yi Zhou 0010, Jiayao Mai, Kuan Dai, Yang Xu 0083, Shaojie Shen |
IEEE Trans. Robotics | 2 |
| 2025 | ESVO2: Direct Visual-Inertial Odometry With Stereo Event CamerasabstractEvent-based visual odometry is a specific branch of visual simultaneous localization and mapping (SLAM) techniques, which aims at solving tracking and mapping subproblems (typically in parallel), by exploiting the special working principles of neuromorphic (i.e., event-based) cameras. Due to the motion-dependent nature of event data, explicit data association (i.e., feature matching) under large-baseline viewpoint changes is difficult to establish, making direct methods a more rational choice. However, state-of-the-art direct methods are limited by the high computational complexity of the mapping subproblem and the degeneracy of camera pose tracking in certain degrees of freedom (DoF) in rotation. In this article, we tackle these issues by building an event-based stereo visual-inertial odometry system, which is built upon a direct pipeline known as event-based stereo visual odometry (ESVO). Specifically, to speed up the mapping operation, we propose an efficient strategy for sampling contour points according to the local dynamics of events. The mapping performance is also improved in terms of structure completeness and local smoothness by merging the temporal stereo and static stereo results. To circumvent the degeneracy of camera pose tracking in recovering the pitch and yaw components of general 6-DoF motion, we introduce IMU measurements as motion priors via preintegration. To this end, a compact back-end is proposed for continuously updating the IMU bias and predicting the linear velocity, enabling an accurate motion prediction for camera pose tracking. The resulting system scales well with modern high-resolution event cameras and leads to better global positioning accuracy in large-scale outdoor environments. Extensive evaluations on five publicly available datasets featuring different resolutions and scenarios justify the superior performance of the proposed system against five state-of-the-art methods. Compared to ESVO, our new pipeline significantly reduces the camera pose tracking error by 40%–80% and 20%–80% in terms of absolute trajectory error and relative pose error, respectively; at the same time, the mapping efficiency is improved by a factor of five. We release our pipeline as an open-source software for future research in this field. Junkai Niu, Xiuyuan Lu, Shaojie Shen, Guillermo Gallego 0002, Yi Zhou 0010 |
IEEE Trans. Robotics | 6 |
| 2024 | Event-Aided Time-to-Collision Estimation for Autonomous Driving
Bangyan Liao, Xiuyuan Lu, Peidong Liu 0001, Shaojie Shen, Yi Zhou 0010 |
ECCV (54) | 6 |
| 2024 | BeNeRF: Neural Radiance Fields from a Single Blurry Image and Event Stream
Wenpu Li, Pian Wan, Peng Wang 0141, Yi Zhou 0010, Peidong Liu 0001 |
ECCV (31) | 5 |
| 2024 | Motion and Structure from Event-Based Normal Flow
Zhongyang Ren, Bangyan Liao, Delei Kong, Peidong Liu 0001, Laurent Kneip, Guillermo Gallego 0002, Yi Zhou 0010 |
ECCV (56) | 8 |
| 2024 | IMU-Aided Event-based Stereo Visual OdometryabstractDirect methods for event-based visual odometry solve the mapping and camera pose tracking sub-problems by establishing implicit data association in a way that the generative model of events is exploited. The main bottlenecks faced by state-of-the-art work in this field include the high computational complexity of mapping and the limited accuracy of tracking. In this paper, we improve our previous direct pipeline Event-based Stereo Visual Odometry in terms of accuracy and efficiency. To speed up the mapping operation, we propose an efficient strategy of edge-pixel sampling according to the local dynamics of events. The mapping performance in terms of completeness and local smoothness is also improved by combining the temporal stereo results and the static stereo results. To circumvent the degeneracy issue of camera pose tracking in recovering the yaw component of general 6-DoF motion, we introduce as a prior the gyroscope measurements via pre-integration. Experiments on publicly available datasets justify our improvement. We release our pipeline as an open-source software for future research in this field. Junkai Niu, Yi Zhou 0010 |
ICRA | 3 |
| 2023 | Event-Based Motion Segmentation With Spatio-Temporal Graph CutsabstractIdentifying independently moving objects is an essential task for dynamic scene understanding. However, traditional cameras used in dynamic scenes may suffer from motion blur or exposure artifacts due to their sampling principle. By contrast, event-based cameras are novel bio-inspired sensors that offer advantages to overcome such limitations. They report pixel-wise intensity changes asynchronously, which enables them to acquire visual information at exactly the same rate as the scene dynamics. We develop a method to identify independently moving objects acquired with an event-based camera, that is, to solve the event-based motion segmentation problem. We cast the problem as an energy minimization one involving the fitting of multiple motion models. We jointly solve two sub-problems, namely event-cluster assignment (labeling) and motion model fitting, in an iterative manner by exploiting the structure of the input event data in the form of a spatio-temporal graph. Experiments on available datasets demonstrate the versatility of the method in scenes with different motion patterns and number of moving objects. The evaluation shows state-of-the-art results without having to predetermine the number of expected moving objects. We release the software and dataset under an open source license to foster research in the emerging topic of event-based motion segmentation. Yi Zhou 0010, Guillermo Gallego 0002, Xiuyuan Lu, Siqi Liu 0022, Shaojie Shen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Event-based Motion Segmentation by Cascaded Two-Level Multi-Model FittingabstractAmong prerequisites for a synthetic agent to inter-act with dynamic scenes, the ability to identify independently moving objects is specifically important. From an application perspective, nevertheless, standard cameras may deteriorate remarkably under aggressive motion and challenging illumination conditions. In contrast, event-based cameras, as a category of novel biologically inspired sensors, deliver advantages to deal with these challenges. Its rapid response and asynchronous nature enables it to capture visual stimuli at exactly the same rate of the scene dynamics. In this paper, we present a cascaded two-level multi-model fitting method for identifying independently moving objects (i.e., the motion segmentation problem) with a monocular event camera. The first level leverages tracking of event features and solves the feature clustering problem under a progressive multi-model fitting scheme. Initialized with the resulting motion model instances, the second level further addresses the event clustering problem using a spatio-temporal graph-cut method. This combination leads to efficient and accurate event-wise motion segmentation that cannot be achieved by any of them alone. Experiments demonstrate the effectiveness and versatility of our method in real-world scenes with different motion patterns and an unknown number of independently moving objects. Xiuyuan Lu, Yi Zhou 0010, Shaojie Shen |
IROS | 2 |
| 2021 | Event-Based Stereo Visual OdometryabstractEvent-based cameras are bioinspired vision sensors whose pixels work independently from each other and respond asynchronously to brightness changes, with microsecond resolution. Their advantages make it possible to tackle challenging scenarios in robotics, such as high-speed and high dynamic range scenes. We present a solution to the problem of visual odometry from the data acquired by a stereo event-based camera rig. Our system follows a parallel tracking-and-mapping approach, where novel solutions to each subproblem (three-dimensional (3-D) reconstruction and camera pose estimation) are developed with two objectives in mind: being principled and efficient, for real-time operation with commodity hardware. To this end, we seek to maximize the spatio-temporal consistency of stereo event-based data while using a simple and efficient representation. Specifically, the mapping module builds a semidense 3-D map of the scene by fusing depth estimates from multiple viewpoints (obtained by spatio-temporal consistency) in a probabilistic fashion. The tracking module recovers the pose of the stereo rig by solving a registration problem that naturally arises due to the chosen map and event data representation. Experiments on publicly available datasets and on our own recordings demonstrate the versatility of the proposed method in natural scenes with general 6-DoF motion. The system successfully leverages the advantages of event-based cameras to perform visual odometry in challenging illumination conditions, such as low-light and high dynamic range, while running in real-time on a standard CPU. We release the software and dataset under an open source license to foster research in the emerging topic of event-based simultaneous localization and mapping. Yi Zhou 0010, Guillermo Gallego 0002, Shaojie Shen |
IEEE Trans. Robotics | 1 |
| 2019 | Canny-VO: Visual Odometry With RGB-D Cameras Based on Geometric 3-D-2-D Edge AlignmentabstractThis paper reviews the classical problem of free-form curve registration and applies it to an efficient RGB-D visual odometry system called Canny-VO, as it efficiently tracks all Canny edge features extracted from the images. Two replacements for the distance transformation commonly used in edge registration are proposed: approximate nearest neighbor fields and oriented nearest neighbor fields. 3-D-2-D edge alignment benefits from these alternative formulations in terms of both efficiency and accuracy. It removes the need for the more computationally demanding paradigms of data-to-model registration, bilinear interpolation, and subgradient computation. To ensure robustness of the system in the presence of outliers and sensor noise, the registration is formulated as a maximum a posteriori problem and the resulting weighted least-squares objective is solved by the iteratively reweighted least-squares method. A variety of robust weight functions are investigated and the optimal choice is made based on the statistics of the residual errors. Efficiency is furthermore boosted by an adaptively sampled definition of the nearest neighbor fields. Extensive evaluations on public SLAM benchmark sequences demonstrate state-of-the-art performance and an advantage over classical Euclidean distance fields. Yi Zhou 0010, Hongdong Li, Laurent Kneip |
IEEE Trans. Robotics | 1 |
| 2018 | Semi-dense 3D Reconstruction with a Stereo Event Camera
Yi Zhou 0010, Guillermo Gallego 0002, Henri Rebecq, Laurent Kneip, Hongdong Li, Davide Scaramuzza 0001 |
ECCV (1) | 1 |
| 2017 | Semi-dense visual odometry for RGB-D cameras using approximate nearest neighbour fieldsabstractThis paper presents a robust and efficient semidense visual odometry solution for RGB-D cameras. The core of our method is a 2D-3D ICP pipeline which estimates the pose of the sensor by registering the projection of a 3D semidense map of a reference frame with the 2D semi-dense region extracted in the current frame. The processing is speeded up by efficiently implemented approximate nearest neighbour fields under the Euclidean distance criterion, which permits the use of compact Gauss-Newton updates in the optimization. The registration is formulated as a maximum a posterior problem to deal with outliers and sensor noise, and the equivalent weighted least squares problem is consequently solved by iteratively reweighted least squares method. A variety of robust weight functions are tested and the optimum is determined based on the probabilistic characteristics of the sensor model. Extensive evaluation on publicly available RGB-D datasets shows that the proposed method predominantly outperforms existing state-of-the-art methods. Yi Zhou 0010, Laurent Kneip, Hongdong Li |
ICRA | 1 |
| 2016 | Divide and Conquer: Efficient Density-Based Tracking of 3D Sensors in Manhattan Worlds
Yi Zhou 0010, Laurent Kneip, Cristian Rodriguez Opazo, Hongdong Li |
ACCV (5) | 1 |
| 2016 | Real-time rotation estimation for dense depth sensors in piece-wise planar environmentsabstractLow-drift rotation estimation is a crucial part of any accurate odometry system. In this paper, we focus on the problem of 3D rotation estimation with dense depth sensors in environments that consist of piece-wise planar structures, such as corridors and office rooms. An efficient mean-shift paradigm is developed to extract and track planar modes in the surface normal vector distribution on the unit sphere. Robust and piece-wise drift-free behavior is achieved by registering the bundle of planar modes from the current frame with respect to a reference frame using a general ℓ1-norm regression scheme. We furthermore add a memory scheme to the regular birth and death of modes, which further compensates accumulated rotational drift when previously discovered modes are revisited. We discuss the robustness issue and evaluate our algorithm on both custom synthetic as well as real publicly available datasets. Our experimental results demonstrate high robustness and effectiveness of the proposed algorithm. Yi Zhou 0010, Laurent Kneip, Hongdong Li |
IROS | 1 |