VLDB 2026 Research / reviewers in the wild / expert
Bogdan Stanciulescu
dblp:11/3564
· DBLP profile ↗
21ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FOOTPASS: A multi-modal multi-agent tactical context dataset for play-by-play action spotting in soccer broadcast videosabstractSoccer video understanding has motivated the creation of datasets for tasks such as temporal action localization, spatiotemporal action detection, jersey recognition, and multi-object tracking. Automating the annotation of structured sequences of events (who does what, when, and where) used in soccer analytics requires a holistic integration of these computer-vision subtasks. However, current action recognition methods remain insufficient for producing reliable play-by-play data and are typically used to assist rather than fully automate annotation. In parallel, research in tactical modeling and trajectory forecasting has highlighted the strong tactical and temporal regularities that structure soccer, motivating the use of tactical knowledge as a prior to support vision-based predictions and enable more automated and reliable play-by-play reconstruction. We introduce Footovision Play-by-Play Action Spotting in Soccer Dataset (FOOTPASS), the first benchmark for play-by-play action spotting over entire soccer matches in a multi-modal, multi-agent tactical context. It supports the development of player-centric action spotting methods that combine computer-vision outputs with long-term tactical regularities inherent to the game. In addition, the dataset benchmarks several spatiotemporal action detection methods, highlighting the limitations of purely perceptual approaches and demonstrating the benefits of incorporating tactical reasoning for improving the reliability and completeness of play-by-play predictions. Jérémie Ochin, Raphaël Chekroun, Bogdan Stanciulescu, Sotiris Manitsaris |
Comput. Vis. Image Underst. | 3 |
| 2025 | Beyond Pixels: Leveraging the Language of Soccer to Improve Spatio-Temporal Action Detection in Broadcast Videos
Jérémie Ochin, Raphaël Chekroun, Bogdan Stanciulescu, Sotiris Manitsaris |
ACIVS | 3 |
| 2025 | Game State and Spatio-Temporal Action Detection in Soccer Using Graph Neural Networks and 3D Convolutional NetworksabstractInternational audience Jérémie Ochin, Guillaume Devineau, Bogdan Stanciulescu, Sotiris Manitsaris |
ICPRAM | 3 |
| 2023 | CROSSFIRE: Camera Relocalization On Self-Supervised Features from an Implicit RepresentationabstractBeyond novel view synthesis, Neural Radiance Fields (NeRF) are useful for applications that interact with the real world. In this paper, we use them as an implicit map of a given scene and propose a camera relocalization algorithm tailored for this representation. The proposed method enables to compute in real-time the precise position of a device using a single RGB camera, during its navigation. In contrast with previous work, we do not rely on pose regression or photometric alignment but rather use dense local features obtained through volumetric rendering which are specialized on the scene with a self-supervised objective. As a result, our algorithm is more accurate than competitors, able to operate in dynamic outdoor environments with changing lightning conditions and can be readily integrated in any volumetric neural renderer. Arthur Moreau, Nathan Piasco, Moussâb Bennehar, Dzmitry Tsishkou, Bogdan Stanciulescu, Arnaud de La Fortelle |
ICCV | 5 |
| 2023 | TSGN: Temporal Scene Graph Neural Networks with Projected Vectorized Representation for Multi-Agent Motion PredictionabstractPredicting future motions of nearby agents is essential for an autonomous vehicle to take safe and effective actions. In this paper, we propose TSGN, a framework using Temporal Scene Graph Neural Networks with projected vectorized representations for multi-agent trajectory prediction. Projected vectorized representation models the traffic scene as a graph which is constructed by a set of vectors. These vectors represent agents, road network, and their spatial relative relationships. All relative features under this representation are both translation-and rotation-invariant. Based on this representation, TSGN captures the spatial-temporal features across agents, road network, interactions among them, and temporal dependencies of temporal traffic scenes. TSGN can predict multimodal future trajectories for all agents simultaneously, plausibly, and accurately. Mean-while, we propose a Hierarchical Lane Transformer for capturing interactions between agents and road network, which filters the surrounding road network and only keeps the most probable lane segments which could have an impact on the future behavior of the target agent. Without sacrificing the prediction performance, this greatly reduces the computational burden. Experiments show TSGN achieves state-of-the-art performance on the Argoverse motion forecasting benchmark. Yunong Wu, Thomas Gilles, Bogdan Stanciulescu, Fabien Moutarde |
IV | 3 |
| 2023 | ImPosing: Implicit Pose Encoding for Efficient Visual LocalizationabstractWe propose a novel learning-based formulation for visual localization of vehicles that can operate in real-time in city-scale environments. Visual localization algorithms determine the position and orientation from which an image has been captured, using a set of geo-referenced images or a 3D scene representation. Our new localization paradigm, named Implicit Pose Encoding (ImPosing), embeds images and camera poses into a common latent representation with 2 separate neural networks, such that we can compute a similarity score for each image-pose pair. By evaluating candidates through the latent space in a hierarchical manner, the camera position and orientation are not directly regressed but incrementally refined. Very large environments force competitors to store gigabytes of map data, whereas our method is very compact independently of the reference database size. In this paper, we describe how to effectively optimize our learned modules, how to combine them to achieve real-time localization, and demonstrate results on diverse large scale scenarios that significantly outperform prior work in accuracy and computational efficiency. Arthur Moreau, Thomas Gilles, Nathan Piasco, Dzmitry Tsishkou, Bogdan Stanciulescu, Arnaud de La Fortelle |
WACV | 5 |
| 2022 | THOMAS: Trajectory Heatmap Output with learned Multi-Agent Sampling
Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, Fabien Moutarde |
ICLR | 4 |
| 2022 | GOHOME: Graph-Oriented Heatmap Output for future Motion EstimationabstractIn this paper, we propose GOHOME, a method leveraging graph representations of the High Definition Map and sparse projections to generate a heatmap output representing the future position probability distribution for a given agent in a traffic scene. This heatmap output yields an unconstrained 2D grid representation of agent future possible locations, allowing inherent multimodality and a measure of the uncertainty of the prediction. Our graph-oriented model avoids the high computation burden of representing the surrounding context as squared images and processing it with classical CNNs, but focuses instead only on the most probable lanes where the agent could end up in the immediate future. GOHOME reaches 2nd on Argoverse Motion Forecasting Benchmark on the Misskate6metric while achieving significant speed-up and memory burden diminution compared to Argoverse 1stplace method HOME. We also highlight that heatmap output enables multimodal ensembling and improve 1stplace MissRate6by more than 15% with our best ensemble on Argoverse. Finally, we evaluate and reach state-of-the-art performance on the other trajectory prediction datasets nuScenes and Interaction, demonstrating the generalizability of our method. Thomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu, Fabien Moutarde |
ICRA | 4 |
| 2022 | Assessing Cross-dataset Generalization of Pedestrian Crossing PredictorsabstractPedestrian crossing prediction has been a topic of active research, resulting in many new algorithmic solutions. While measuring the overall progress of those solutions over time tends to be more and more established due to the new publicly available benchmark and standardized evaluation procedures, knowing how well existing predictors react to unseen data remains an unanswered question. This evaluation is imperative as serviceable crossing behavior predictors should be set to work in various scenarios without compromising pedestrian safety due to misprediction. To this end, we conduct a study based on direct cross-dataset evaluation. Our experiments show that current state-of-the-art pedestrian behavior predictors generalize poorly in cross-dataset evaluation scenarios, regardless of their robustness during a direct training-test set evaluation setting. In the light of what we observe, we argue that the future of pedestrian crossing prediction, e.g. reliable and generalizable implementations, should not be about tailoring models, trained with very little available data, and tested in a classical train-test scenario with the will to infer anything about their behavior in real life. It should be about evaluating models in a cross-dataset setting while considering their uncertainty estimates under domain shift. Joseph Gesnouin, Steve Pechberti, Bogdan Stanciulescu, Fabien Moutarde |
IV | 3 |
| 2022 | CoordiNet: uncertainty-aware pose regressor for reliable vehicle localizationabstractIn this paper, we investigate visual-based camera re-localization with neural networks for robotics and autonomous vehicles applications. Our solution is a CNN-based algorithm which predicts camera pose (3D translation and 3D rotation) directly from a single image. It also provides an uncertainty estimate of the pose. Pose and uncertainty are learned together with a single loss function and are fused at test time with an EKF. Furthermore, we propose a new fully convolutional architecture, named CoordiNet, designed to embed some of the scene geometry.Our framework outperforms comparable methods on the largest available benchmark, the Oxford RobotCar dataset, with an average error of 8 meters where previous best was 19 meters. We have also investigated the performance of our method on large scenes for real time (18 fps) vehicle localization. In this setup, structure-based methods require a large database, and we show that our proposal is a reliable alternative, achieving 29cm median error in a 1.9km loop in a busy urban area. Arthur Moreau, Nathan Piasco, Dzmitry Tsishkou, Bogdan Stanciulescu, Arnaud de La Fortelle |
WACV | 4 |
| 2021 | TrouSPI-Net: Spatio-temporal attention on parallel atrous convolutions and U-GRUs for skeletal pedestrian crossing predictionabstractUnderstanding the behaviors and intentions of pedestrians is still one of the main challenges for vehicle autonomy, as accurate predictions of their intentions can guarantee their safety and driving comfort of vehicles. In this paper, we address pedestrian crossing prediction in urban traffic environments by linking the dynamics of a pedestrian's skeleton to a binary crossing intention. We introduce TrouSPI-Net: a context-free, lightweight, multi-branch predictor. TrouSPI-Net extracts spatio-temporal features for different time resolutions by encoding pseudo-images sequences of skeletal joints' positions and processes them with parallel attention modules and atrous convolutions. The proposed approach is then enhanced by processing features such as relative distances of skeletal joints, bounding box positions, or ego-vehicle speed with U-GRUs. Using the newly proposed evaluation procedures for two large public naturalistic data sets for studying pedestrian behavior in traffic: JAAD and PIE, we evaluate TrouSPI-Net and analyze its performance. Experimental results show that TrouSPI-Net achieved 76% F1 score on JAAD and 80% F1 score on PIE, therefore outperforming current state-of-the-art while being lightweight and context-free. Joseph Gesnouin, Steve Pechberti, Bogdan Stanciulescu, Fabien Moutarde |
FG | 3 |
| 2017 | Real-time method for general road segmentationabstractImage road detection in unstructured environments is a crucial and challenging problem in the application of mobile robots and autonomous vehicles. In this paper, we present an effective and computationally efficient solution to segment the road region for structured and unstructured roads. We propose a new method that incorporates two different approaches: road detection based on the vanishing point and image segmentation using a seeded region growing (SRG) algorithm. First, a fast vanishing point detection algorithm is applied and used to find an estimation of the road boundaries. Subsequently, we segment the road area region executing a SRG algorithm based on the vanishing point and the road boundaries found previously. Evaluation of our method over different images datasets demonstrates that it is effective in challenging conditions such as dirt and curved roads. Michelle Valente, Bogdan Stanciulescu |
Intelligent Vehicles Symposium | 2 |
| 2015 | Clustering and classifying deformations for Shape Regression applied to the human bodyabstractIn this paper, we present an approach to the problem of aligning a Point Distribution Model onto the human body. The key idea of this paper is the clustering of the shape displacements into classes, in order to group the residual deformations in a coherent direction according to the image. We combine this approach by employing the implicitly encoded constraints and the double cascade pattern defined in the “Explicit Shape Regression”. We proceed with adjustments adapted for the context of human body. Hence, we reformulate a regression problem into a classification of displacements. We show through our experiments that obtaining minimal classification errors leads effectively to faster and more accurate regression results. We enhance this classification process by conjointly transforming residual displacements and images to train our classifiers. Olivier Huynh, Bogdan Stanciulescu |
ISPA | 2 |
| 2015 | Person Re-identification Using the Silhouette Shape Described by a Point Distribution ModelabstractIn this paper, we present a new shape-based system for person re-identification. The silhouette shape is represented by a Point Distribution Model (PDM) aligned on the body. We improve a fitting model which iteratively adjusts the shape by maximizing a boosted score of local features: the "Boosted Deformable Model". We modify the training procedure with a ranking structure to find how the model can approach the correct fitting. This is enhanced by the use of weak Artificial Neural Networks as regression functions. Then, we experiment the use of two kind of descriptors on the aligned model : a pose shape signature built with the Shape Context on the set of landmarks and an appearance-based signature using color histograms on the warped appearance contained in the shape model. We demonstrate our approach with evaluations employing the alignment and re-identification modules. The results show that our improvements provide a more accurate fitting, the adapted shape representation has a potential discriminant for re-identification through the pose and employing a PDM instead of a pixel mask to describe silhouette enhances performance of conventional appearance features. Olivier Huynh, Bogdan Stanciulescu |
WACV | 2 |
| 2014 | Hybrid visual and inertial RANSAC for real-time motion estimationabstractWe present a real-time, online motion estimation algorithm combining inertial and visual measurements. The approach is built upon the preemptive RANSAC, a real-time motion estimation algorithm. Our algorithm extends the preemptive RANSAC with inertial sensors data, introducing a lagrangian hybrid scoring of the motion models. We also improve the system using a pure inertial model that is scored as the visual ones, and a dynamic computation of the lagrangian to make the approach adaptive to various image and motion contents. The algorithm is run frame to frame to avoid error accumulation. All these improvements are made with little computational cost, keeping the complexity of the algorithm low enough for embedded platforms. The approach is compared with pure inertial and pure visual procedures. Manu Alibay, Stephane Auberger, Bogdan Stanciulescu, Philippe Fuchs |
ICIP | 3 |
| 2014 | General Road Detection Algorithm - A Computational ImprovementabstractInternational audience Bruno Ricaud, Bogdan Stanciulescu, Amaury Breheret |
ICPRAM | 2 |
| 2012 | Real-Time Traffic-Sign Recognition Using Tree ClassifiersabstractTraffic-sign recognition (TSR) is an essential component of a driver assistance system (DAS), providing drivers with safety and precaution information. In this paper, we evaluate the performance of k-d trees, random forests, and support vector machines (SVMs) for traffic-sign classification using different-sized histogram-of-oriented-gradient (HOG) descriptors and distance transforms (DTs). We also use the Fisher's criterion and random forests for the feature selection to reduce the memory requirements and enhance the performance. We use the German Traffic Sign Recognition Benchmark (GTSRB) data set containing 43 classes and more than 50 000 images. Fatin Zaklouta, Bogdan Stanciulescu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2011 | Traffic sign classification using K-d trees and Random ForestsabstractIn this paper, we evaluate the performance of K-d trees and Random Forests for traffic sign classification using different size Histogram of Oriented Gradients (HOG) descriptors and Distance Transforms. We use the German Traffic Sign Benchmark data set [1] containing 43 classes and more than 50,000 images. The K-d tree is fast to build and search in. We combine the tree classifiers with the HOG descriptors as well as the Distance Transforms and achieve classification rates of up to 97% and 81.8% respectively. Fatin Zaklouta, Bogdan Stanciulescu, Omar Hamdoun |
IJCNN | 2 |
| 2011 | Warning traffic sign recognition using a HOG-based K-d treeabstractTraffic signs are an essential part of a Driver Assistance System (DAS) and provide drivers with safety information. They are designed to be easily seen and understood. The triangular signs warn the drivers of imminent dangers such as wild animals or a sharp curve. In this paper, an efficient algorithm for the detection and recognition of warning signs is presented. A Histogram of Oriented Gradients (HOG) is used to detect 95% of the triangular warning signs. A blackhat filter eliminates a large part of the false alarms. An approximate nearest neighbors search using a KD-tree refines the result. It eliminates 100% of the remaining false detections and distinguishes amongst the different types of signs. The advantage of using HOG features is that all the warning signs, including static (red frame) and dynamic warning signs (illuminated) can be detected with a single detector and therefore, only one image scan. Fatin Zaklouta, Bogdan Stanciulescu |
Intelligent Vehicles Symposium | 2 |
| 2010 | Classifying bags of keypoints using HMMsabstractIn this paper, we use a Hidden Markov Models (HMM) to classify bags of SURF keypoints descriptors of a given class. The performance of this technique is compared to that of others, by testing it on various multi-class datasets. We also describe a prospective of expanding our application to include the detection and classification of moving objects in a video stream using optical flow and Self Organizing Maps (SOM). Fatin Zaklouta, Bogdan Stanciulescu |
AICCSA | 2 |
| 2003 | Physical modeling framework for robotics applicationsabstractThis paper shows how it is possible to extend a discrete physics modeling framework, the mass-springs systems used to describe objects in animation, to robotics applications. Previous work on mass-springs systems resulted in modeling physical phenomena or in artificial life animation synthesis. This paper analyses the possibility to control a physical system by another in order to model complex heterogeneous time-dependent systems, as found in robotics applications or in artificial life. The work described is based on Cordis-Anima, a particle-based modeling framework. An application of a mobile-following-target, in which both mobile and controller are complex mass-springs structures and all interaction forces with the environment are simulated, is presented. Several synthesis animation frames are presented showing the mobile and controller's behavior. Bogdan Stanciulescu, Jean-Loup Florens, Annie Luciani, Jean Louchet |
SMC | 1 |