EDBT 2026 Demo / reviewers in the wild / expert
Dzmitry Tsishkou
dblp:40/5636
· DBLP profile ↗
15ranked-venue papers
0as first author
13since 2021 · last 2025
0009-0002-9798-3316ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Systems, architecture and hardware · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ViiNeuS: Volumetric Initialization for Implicit Neural Surface Reconstruction of Urban Scenes with Limited Image OverlapabstractNeural implicit surface representation methods have recently shown impressive 3D reconstruction results. However, existing solutions struggle to reconstruct driving scenes due to their large size, highly complex nature and their limited visual observation overlap. Hence, to achieve accurate reconstructions, additional supervision data such as LiDAR, strong geometric priors, and long training times are required. To tackle such limitations, we present ViiNeuS, a new hybrid implicit surface learning method that efficiently initializes the signed distance field to reconstruct large driving scenes from 2D street view images. ViiNeuS’s hybrid architecture models two separate implicit fields: one representing the volumetric density of the scene, and another one representing the signed distance to the surface. To accurately reconstruct urban outdoor driving scenarios, we introduce a novel volume-rendering strategy that relies on self-supervised probabilistic density estimation to sample points near the surface and transition progressively from volumetric to surface representation. Our solution permits a proper and fast initialization of the signed distance field without relying on any geometric prior on the scene, compared to concurrent methods. By conducting extensive experiments on four outdoor driving datasets, we show that ViiNeuS can learn an accurate and detailed 3D surface representation of various urban scene while being two times faster to train compared to previous state-of-the-art solutions. Hala Djeghim, Nathan Piasco, Moussâb Bennehar, Luis Roldão, Dzmitry Tsishkou, Desire Sidibé |
CVPR | 5 |
| 2024 | PlaNeRF: SVD Unsupervised 3D Plane Regularization for NeRF Large-Scale Urban Scene ReconstructionabstractNeural Radiance Fields (NeRF) enable 3D scene reconstruction from 2D images and camera poses for Novel View Synthesis (NVS). Although NeRF can produce photorealistic results, it often suffers from overfitting to training views, leading to poor geometry reconstruction, especially in low-texture areas such as road surfaces in driving scenarios. This limitation restricts many important applications which require accurate geometry, such as extrapolated NVS, HD mapping, simulation and scene editing. To address this limitation, we propose a new method to improve NeRF’s 3D structure using only RGB images and semantic maps. Our approach introduces a novel plane regularization based on Singular Value Decomposition (SVD), that does not rely on any geometric prior. In addition, we leverage the Structural Similarity Index Measure (SSIM) in patch-based loss design to properly initialize the volumetric representation of NeRF. Quantitative and qualitative results show that our method outperforms popular regularization approaches in accurate geometry reconstruction for large-scale outdoor scenes and achieves comparable rendering quality to SOTA methods on the KITTI-360 NVS benchmark. Fusang Wang, Arnaud Louys, Nathan Piasco, Moussâb Bennehar, Luis Roldão, Dzmitry Tsishkou |
3DV | 6 |
| 2024 | SOAC: Spatio-Temporal Overlap-Aware Multi-Sensor Calibration using Neural Radiance FieldsabstractIn rapidly-evolving domains such as autonomous driving, the use of multiple sensors with different modalities is crucial to ensure high operational precision and stability. To correctly exploit the provided information by each sensor in a single common frame, it is essential for these sensors to be accurately calibrated. In this paper, we leverage the ability of Neural Radiance Fields (NeRF) to represent different sensors modalities in a common volumetric representation to achieve robust and accurate spatio-temporal sensor calibration. By designing a partitioning approach based on the visible part of the scene for each sensor, we formulate the calibration problem using only the overlapping areas. This strategy results in a more robust and accurate calibration that is less prone to failure. We demonstrate that our approach works on outdoor urban scenes by validating it on multiple established driving datasets. Results show that our method is able to get better accuracy and robustness compared to existing methods. Quentin Herau, Nathan Piasco, Moussâb Bennehar, Luis Roldão, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, Cédric Demonceaux |
CVPR | 5 |
| 2024 | SWAG: Splatting in the Wild Images with Appearance-Conditioned Gaussians
Hiba Dahmani, Moussâb Bennehar, Nathan Piasco, Luis Roldão, Dzmitry Tsishkou |
ECCV (76) | 5 |
| 2024 | RoDUS: Robust Decomposition of Static and Dynamic Elements in Urban Scenes
Thang-Anh-Quan Nguyen, Luis Roldão, Nathan Piasco, Moussâb Bennehar, Dzmitry Tsishkou |
ECCV (72) | 5 |
| 2024 | 3DGS-Calib: 3D Gaussian Splatting for Multimodal SpatioTemporal CalibrationabstractReliable multimodal sensor fusion algorithms require accurate spatiotemporal calibration. Recently, targetless calibration techniques based on implicit neural representations have proven to provide precise and robust results. Nevertheless, such methods are inherently slow to train given the high computational overhead caused by the large number of sampled points required for volume rendering. With the recent introduction of 3D Gaussian Splatting as a faster alternative to implicit representation methods, we propose to leverage this new rendering approach to achieve faster multi-sensor calibration. We introduce 3DGS-Calib, a new calibration method that relies on the speed and rendering accuracy of 3D Gaussian Splatting to achieve multimodal spatiotemporal calibration that is accurate, robust, and with a substantial speed-up compared to methods relying on implicit neural representations. We demonstrate the superiority of our proposal with experimental results on sequences from KITTI-360, a widely used driving dataset. Quentin Herau, Moussâb Bennehar, Arthur Moreau, Nathan Piasco, Luis Roldão, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, Cédric Demonceaux |
IROS | 6 |
| 2023 | CROSSFIRE: Camera Relocalization On Self-Supervised Features from an Implicit RepresentationabstractBeyond novel view synthesis, Neural Radiance Fields (NeRF) are useful for applications that interact with the real world. In this paper, we use them as an implicit map of a given scene and propose a camera relocalization algorithm tailored for this representation. The proposed method enables to compute in real-time the precise position of a device using a single RGB camera, during its navigation. In contrast with previous work, we do not rely on pose regression or photometric alignment but rather use dense local features obtained through volumetric rendering which are specialized on the scene with a self-supervised objective. As a result, our algorithm is more accurate than competitors, able to operate in dynamic outdoor environments with changing lightning conditions and can be readily integrated in any volumetric neural renderer. Arthur Moreau, Nathan Piasco, Moussâb Bennehar, Dzmitry Tsishkou, Bogdan Stanciulescu, Arnaud de La Fortelle |
ICCV | 4 |
| 2023 | MOISST: Multimodal Optimization of Implicit Scene for SpatioTemporal CalibrationabstractWith the recent advances in autonomous driving and the decreasing cost of LiDARs, the use of multimodal sensor systems is on the rise. However, in order to make use of the information provided by a variety of complimentary sensors, it is necessary to accurately calibrate them. We take advantage of recent advances in computer graphics and implicit volumetric scene representation to tackle the problem of multi-sensor spatial and temporal calibration. Thanks to a new formulation of the Neural Radiance Field (NeRF) optimization, we are able to jointly optimize calibration parameters along with scene representation based on radiometric and geometric measurements. Our method enables accurate and robust calibration from data captured in uncontrolled and unstructured urban environments, making our solution more scalable than existing calibration solutions. We demonstrate the accuracy and robustness of our method in urban scenes typically encountered in autonomous driving scenarios. Quentin Herau, Nathan Piasco, Moussâb Bennehar, Luis Roldão, Dzmitry Tsishkou, Cyrille Migniot, Pascal Vasseur, Cédric Demonceaux |
IROS | 5 |
| 2023 | ImPosing: Implicit Pose Encoding for Efficient Visual LocalizationabstractWe propose a novel learning-based formulation for visual localization of vehicles that can operate in real-time in city-scale environments. Visual localization algorithms determine the position and orientation from which an image has been captured, using a set of geo-referenced images or a 3D scene representation. Our new localization paradigm, named Implicit Pose Encoding (ImPosing), embeds images and camera poses into a common latent representation with 2 separate neural networks, such that we can compute a similarity score for each image-pose pair. By evaluating candidates through the latent space in a hierarchical manner, the camera position and orientation are not directly regressed but incrementally refined. Very large environments force competitors to store gigabytes of map data, whereas our method is very compact independently of the reference database size. In this paper, we describe how to effectively optimize our learned modules, how to combine them to achieve real-time localization, and demonstrate results on diverse large scale scenarios that significantly outperform prior work in accuracy and computational efficiency. Arthur Moreau, Thomas Gilles, Nathan Piasco, Dzmitry Tsishkou, Bogdan Stanciulescu, Arnaud de La Fortelle |
WACV | 4 |
| 2022 | Learning Human-like Driving Policies from Real Interactive Driving ScenesabstractApprentissage de politiques de navigation a partir de demonstration pour de la simulation de traffic Yann Koeberle, Stefano Sabatini, Dzmitry Tsishkou, Christophe Sabourin |
ICINCO | 3 |
| 2022 | THOMAS: Trajectory Heatmap Output with learned Multi-Agent Sampling
Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, Fabien Moutarde |
ICLR | 3 |
| 2022 | GOHOME: Graph-Oriented Heatmap Output for future Motion EstimationabstractIn this paper, we propose GOHOME, a method leveraging graph representations of the High Definition Map and sparse projections to generate a heatmap output representing the future position probability distribution for a given agent in a traffic scene. This heatmap output yields an unconstrained 2D grid representation of agent future possible locations, allowing inherent multimodality and a measure of the uncertainty of the prediction. Our graph-oriented model avoids the high computation burden of representing the surrounding context as squared images and processing it with classical CNNs, but focuses instead only on the most probable lanes where the agent could end up in the immediate future. GOHOME reaches 2nd on Argoverse Motion Forecasting Benchmark on the Misskate6metric while achieving significant speed-up and memory burden diminution compared to Argoverse 1stplace method HOME. We also highlight that heatmap output enables multimodal ensembling and improve 1stplace MissRate6by more than 15% with our best ensemble on Argoverse. Finally, we evaluate and reach state-of-the-art performance on the other trajectory prediction datasets nuScenes and Interaction, demonstrating the generalizability of our method. Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, Fabien Moutarde |
ICRA | 3 |
| 2022 | CoordiNet: uncertainty-aware pose regressor for reliable vehicle localizationabstractIn this paper, we investigate visual-based camera re-localization with neural networks for robotics and autonomous vehicles applications. Our solution is a CNN-based algorithm which predicts camera pose (3D translation and 3D rotation) directly from a single image. It also provides an uncertainty estimate of the pose. Pose and uncertainty are learned together with a single loss function and are fused at test time with an EKF. Furthermore, we propose a new fully convolutional architecture, named CoordiNet, designed to embed some of the scene geometry.Our framework outperforms comparable methods on the largest available benchmark, the Oxford RobotCar dataset, with an average error of 8 meters where previous best was 19 meters. We have also investigated the performance of our method on large scenes for real time (18 fps) vehicle localization. In this setup, structure-based methods require a large database, and we show that our proposal is a reliable alternative, achieving 29cm median error in a 1.9km loop in a busy urban area. Arthur Moreau, Nathan Piasco, Dzmitry Tsishkou, Bogdan Stanciulescu, Arnaud de La Fortelle |
WACV | 3 |
| 2017 | Efficient combination of Lidar intensity and 3D information by DNN for pedestrian recognition with high and low density 3D sensorabstractPedestrian recognition is one of the key components for assisted and autonomous driving. So far many researchers have investigated systems combining a high density LIDAR with cameras or stereo, which results in an expensive and complex setup where the LIDAR data is mostly used to extract regions of interest for the 2D sensor. Very few work has focused on using pure 3D data coming from the LIDAR to recognize pedestrians, and even less have made an intensive use of the intensity information returned by the LIDAR. The intensity information displays a high frequency change between neighboring points of similar material, this can be due to the angle or distance. Due to this, it has not been frequently investigated as a potentially interesting feature as it would require extensive time consuming feature engineering to be worthwhile. In this paper we present a novel 2D representation of a 3D point cloud including the intensity information. We show the ability of convolutional neural networks to handle this data in order to accurately recognize pedestrians in complex driving scenes. Our system outperformed state of the art technique on the STC database. Additionally we show that this system is still highly accurate on low density LIDAR data. Luc Mioulet, Dzmitry Tsishkou, Rémy Bendahan, Frédéric Abad |
Intelligent Vehicles Symposium | 2 |
| 2008 | Monocular vision obstacles detection for autonomous navigationabstractAutonomous robot navigation has many applications such as space exploration and autonomous vehicles. Currently, such navigation is ensured by the use of multiple sensors which may hinder quick commercialization. In this article, we propose a solution for navigating a robot in an unknown environment using only monocular vision algorithms. S. Wybo, Dzmitry Tsishkou, C. Vestri, Frédéric Abad, S. Bougnoux, Rémy Bendahan |
IROS | 2 |