Jérôme Royan

dblp:42/5562 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0001-9485-8638ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%
Computer graphics and multimedia
2 papers
Visualization and visual analytics · 57% Virtual and augmented reality · 26% Rendering · 17%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Collaborative and social computing › remote collaboration
asymmetric collaboration
0.212015
Laying out spaces with virtual reality · VR 2015
Collaborative and social computing › collaborative virtual environments › social virtual reality
virtual reality collaboration
0.212015
Laying out spaces with virtual reality · VR 2015
Visualization and visual analytics › visual attention
visual attention model
0.112012
Design and Application of Real-Time Visual Attention Model for the Exploration of 3D Virtual Environments · IEEE Trans. Vis. Comput. Graph. 2012
Rendering
level-of-detail rendering
0.012012
Design and Application of Real-Time Visual Attention Model for the Exploration of 3D Virtual Environments · IEEE Trans. Vis. Comput. Graph. 2012

Methods — techniques the papers use, named apart from their topics

asymmetric interaction design · 0.4surface-element representation · 0.1bottom-up and top-down attention · 0.1
YearPublicationVenuePosition
2026 Real-Time Retrieval-Free Camera Pose Estimation via Sparse Cross-Modal 2D-3D Matching with Projection-Guided Refinement
abstract
Camera pose estimation is a fundamental task in visual localization but often relies on image retrieval or dense rendering, leading to high memory and computational cost. We propose a real-time retrieval-free camera pose estimation framework based on sparse cross-modal 2D–3D matching. Our approach learns descriptors for a compact point cloud using dense-render guidance and performs localization through projection-guided uncertainty-aware correspondence refinement, avoiding image database search and dense matching.
Amine Kacete, Jérôme Royan, Guillaume Moreau
ICMR3
2023 TwistSLAM++: Fusing Multiple Modalities for Accurate Dynamic Semantic SLAM
abstract
Most classical SLAM systems rely on the static scene assumption, which limits their applicability in real world scenarios. Recent SLAM frameworks have been proposed to simultaneously track the camera and moving objects. However they are often unable to estimate the canonical pose of the objects and exhibit a low object tracking accuracy. To solve this problem we propose TwistSLAM++, a semantic, dynamic, SLAM system that fuses stereo images and LiDAR information. Using semantic information, we track potentially moving objects and associate them to 3D object detections in LiDAR scans to obtain their pose and size. Then, we perform registration on consecutive object scans to refine object pose estimation. Finally, object scans are used to estimate the shape of the object and constrain map points to lie on the estimated surface within the bundle adjustment. We show on classical benchmarks that this fusion approach based on multimodal information improves the accuracy of object tracking.
Mathieu Gonzalez, Éric Marchand, Amine Kacete, Jérôme Royan
IROS4
2023 MES-Loss: Mutually equidistant separation metric learning loss function
Yasser Boutaleb, Catherine Soladié, Nam-Duong Duong, Amine Kacete, Jérôme Royan, Renaud Séguier
Pattern Recognit. Lett.5
2022 S3LAM: Structured Scene SLAM
abstract
We propose a new SLAM system that uses the semantic segmentation of objects and structures in the scene. Semantic information is relevant as it contains high level information which may make SLAM more accurate and robust. Our contribution is twofold: i) A new SLAM system based on ORB-SLAM2 that creates a semantic map made of clusters of points corresponding to objects instances and structures in the scene. ii) A modification of the classical Bundle Adjustment formulation to constrain each cluster using geometrical priors, which improves both camera localization and reconstruction and enables a better understanding of the scene. We evaluate our approach on sequences from several public datasets and show that it improves camera pose estimation with respect to state of the art.
Mathieu Gonzalez, Éric Marchand, Amine Kacete, Jérôme Royan
IROS4
2020 Efficient multi-output scene coordinate prediction for fast and accurate camera relocalization from a single RGB image
Nam-Duong Duong, Catherine Soladié, Amine Kacete, Pierre-Yves Richard, Jérôme Royan
Comput. Vis. Image Underst.5
2018 Accurate Sparse Feature Regression Forest Learning for Real-Time Camera Relocalization
abstract
Camera relocalization is needed in several applications such as augmented reality or robot navigation. However, it is still challenging to have a both real-time and accurate method. In this paper, we present our hybrid method combing machine learning approach and geometric approach for real-time camera relocalization from a single RGB image. We introduce our sparse feature regression forest to improve the machine learning part. In our regression forest, we propose a novel split function, that uses a whole feature vector instead of classical binary test function to improve the accuracy of 2D-3D point correspondences. Moreover, we use sparse feature extraction (SURF features) to reduce time processing. The results indicate that our method is the only real-time hybrid method (50ms per frame). We also achieve results as accurate as the best state-of-the-art methods (hybrid methods) and outperform machine learning based and sparse feature based methods.
Nam-Duong Duong, Amine Kacete, Catherine Soladié, Pierre-Yves Richard, Jérôme Royan
3DV5
2017 Collaborators awareness for user cohabitation in co-located collaborative virtual environments
abstract
In a co-located collaborative virtual environment, multiple users share the same physical tracked space and the same virtual workspace. When the virtual workspace is larger than the real workspace, navigation interaction techniques must be deployed to let the users explore the entire virtual environment. When a user navigates in the virtual space while remaining static in the real space, his/her position in the physical workspace and in the virtual workspace are no longer the same. Thus, in the context where each user is immersed in the virtual environment with a Head-Mounted-Display, a user can still perceive where his/her collaborators are in the virtual environment but not where they are in real world. In this paper, we propose and compare three methods to warn users about the position of collaborators in the shared physical workspace to ensure a proper cohabitation and safety of the collaborators. The frst one is based on a virtual grid shaped as a cylinder, the second one is based on a ghost representation of the user and the last one displays the physical safe-navigation space on the foor of the virtual environment. We conducted a user-study with two users wearing a Head-Mounted-Display in the context of a collaborative First-Person-Shooter game. Our three methods were compared with a condition where the physical tracked space was separated into two zones, one per user, to evaluate the impact of each condition on safety, displacement freedom and global satisfaction of users. Results suggest that the ghost avatar and the cylinder grid can be good alternatives to the separation of the tracked space.
Jérémy Lacoche, Nico Pallamin, Thomas Boggini, Jérôme Royan
VRST4
2016 Unconstrained Gaze Estimation Using Random Forest Regression Voting
Amine Kacete, Renaud Séguier, Michel Collobert, Jérôme Royan
ACCV (3)4
2016 Real-time eye pupil localization using Hough regression forest
abstract
Eyes are one of the most salient features of the human face, and the location of the pupil allows access to important information which can be used in several computer vision applications. Several commercial eye-trackers can estimate with good accuracy the pupil location, but need complex hardware specifications and a controlled user environment (high eye image resolution, good illumination, small head pose variations) making these solutions difficult to use in an arbitrary environment. In this paper, we present an approach based on Hough randomized regression trees. We demonstrate, by several evaluations on challenging public datasets that our approach is very robust to illumination, scale, eye movements and high head pose variations and yields a significant improvement compared to a wide range of state-of-the-art methods.
Amine Kacete, Jérôme Royan, Renaud Séguier, Michel Collobert, Catherine Soladié
WACV2
2015 Laying out spaces with virtual reality
abstract
When dealing with real estate business, it is quite difficult for estate agents to make customers understand the potential and the volumes of free spaces. Thus, we propose an application that aims to solve these issues based on a laying out scenario in which a seller and a customer collaborate. As the roles of both users are different, we propose an asymmetric collaboration where the two users do not use the same interaction setup and do not benefit from the same interaction capabilities.
Morgan Le Chénéchal, Jérémy Lacoche, Cyndie Martin, Jérôme Royan
VR4
2012 Design and Application of Real-Time Visual Attention Model for the Exploration of 3D Virtual Environments
abstract
This paper studies the design and application of a novel visual attention model designed to compute user's gaze position automatically, i.e., without using a gaze-tracking system. The model we propose is specifically designed for real-time first-person exploration of 3D virtual environments. It is the first model adapted to this context which can compute in real time a continuous gaze point position instead of a set of 3D objects potentially observed by the user. To do so, contrary to previous models which use a mesh-based representation of visual objects, we introduce a representation based on surface-elements. Our model also simulates visual reflexes and the cognitive processes which take place in the brain such as the gaze behavior associated to first-person navigation in the virtual environment. Our visual attention model combines both bottom-up and top-down components to compute a continuous gaze point position on screen that hopefully matches the user's one. We conducted an experiment to study and compare the performance of our method with a state-of-the-art approach. Our results are found significantly better with sometimes more than 100 percent of accuracy gained. This suggests that computing a gaze point in a 3D virtual environment in real time is possible and is a valid approach, compared to object-based approaches. Finally, we expose different applications of our model when exploring virtual environments. We present different algorithms which can improve or adapt the visual feedback of virtual environments based on gaze information. We first propose a level-of-detail approach that heavily relies on multiple-texture sampling. We show that it is possible to use the gaze information of our visual attention model to increase visual quality where the user is looking, while maintaining a high-refresh rate. Second, we introduce the use of the visual attention model in three visual effects inspired by the human visual system namely: depth-of-field blur, camera- motions, and dynamic luminance. All these effects are computed based on the simulated gaze of the user, and are meant to improve user's sensations in future virtual reality applications.
Sébastien Hillaire, Anatole Lécuyer, Tony Regia-Corte, Rémi Cozot, Jérôme Royan, Gaspard Breton
IEEE Trans. Vis. Comput. Graph.5
2010 A real-time visual attention model for predicting gaze point during first-person exploration of virtual environments
abstract
This paper introduces a novel visual attention model to compute user's gaze position automatically, i.e. without using a gaze-tracking system. Our model is specifically designed for real-time first-person exploration of 3D virtual environments. It is the first model adapted to this context which can compute, in real-time, a continuous gaze point position instead of a set of 3D objects potentially observed by the user. To do so, contrary to previous models which use a mesh-based representation of visual objects, we introduce a representation based on surface-elements. Our model also simulates visual reflexes and the cognitive process which takes place in the brain such as the gaze behavior associated to first-person navigation in the virtual environment. Our visual attention model combines the bottom-up and top-down components to compute a continuous gaze point position on screen that hopefully matches the user's one. We have conducted an experiment to study and compare the performance of our method with a state-of-the-art approach. Our results are found significantly better with more than 100% of accuracy gained. This suggests that computing in realtime a gaze point in a 3D virtual environment is possible and is a valid approach as compared to object-based approaches.
Sébastien Hillaire, Anatole Lécuyer, Tony Regia-Corte, Rémi Cozot, Jérôme Royan, Gaspard Breton
VRST5
2009 Peer-to-peer visualization of very large 3D landscape and city models using MPEG-4
Romain Cavagna, Jérôme Royan, Patrick Gioia, Christian Bouville, Maha Abdallah, Eliya Buyukkaya
Signal Process. Image Commun.2
2008 A MPEG-4 AFX compliant platform for 3D contents distribution in peer-to-peer
abstract
Although compression of 3D content has been a widely studied topic, adaptive streaming of large virtual scenes is progressively becoming the standard way to share collaborative environments, due to the growing complexity of the shared data. The classical 'download and play' paradigm associated to standalone compression is leaving place to interactive refinements of regions-of-interest, provided given polygon budget and network capability. With the increasing number of applications and potential clients sharing reusable data, the traditional client / server architecture is likely to be over passed by less server-centered protocols, in order to distribute the complex tasks of identifying the selected data to be streamed. In this paper, we present the design and performances of a platform exploiting peer-to-peer transmission and show how the whole setting can be implemented using the MPEG-4 AFX standards.
Romain Cavagna, Christian Bouville, Patrick Gioia, Jérôme Royan
ICIP4
2006 Real Time P2P Network Simulation for Very Large Virtual Environment
abstract
The ever increasing speed of Internet connections has led to a point where it is actually possible for every end user to seamlessly share data on Internet. Peer-to-peer (P2P) networks are typical of this evolution. The goal of our paper is to show that thanks to self-adaptive assignment techniques, server-less P2P networks can efficiently deal with very large environments such as met in the geo-visualisation area. Our method takes advantage of a hierarchical and progressive data structure that describes the environment. In order to assess the global efficiency of this P2P technique, we have implemented a dedicated real time simulator. Experimentation results are presented using a hierarchical LOD model of a very large urban environment
Romain Cavagna, Christian Bouville, Jérôme Royan
DS-RT3
2006 P2P Network for very large virtual environment
abstract
The ever increasing speed of Internet connections has led to a point where it is actually possible for every end user to seamlessly share data on Internet. Peer-To-Peer (P2P) networks are typical of this evolution. The goal of our paper is to show that server-less P2P networks with self-adaptive assignment techniques can efficiently deal with very large environments such as met in the geovisualization domain. Our method allows adaptative view-dependent visualization thanks to a hierarchical and progressive data structure that describes the environment. In order to assess the global efficiency of this P2P technique, we have implemented a dedicated real time simulator. Experimentation results are presented using a hierarchical LOD model of a very large urban environment.
Romain Cavagna, Christian Bouville, Jérôme Royan
VRST3