Jordi Sanchez-Riera

dblp:73/1445 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 68% Vision and language · 32%
Human-computer interaction and pervasive computing
1 paper
Accessibility and assistive technology · 77% Wearable and physiological sensing · 23%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
visual question answering
0.912025
VQA-Driven Event Maps for Assistive Navigation for People with Low Vision in Urban Environments · ICRA 2025
Accessibility and assistive technology
assistive navigation
0.912025
VQA-Driven Event Maps for Assistive Navigation for People with Low Vision in Urban Environments · ICRA 2025
Computer vision › 3D vision › motion estimation › non-rigid motion estimation
human motion estimation
0.812024
MultiPhys: Multi-Person Physics-Aware 3D Motion Estimation · CVPR 2024
Computer vision › 3D vision › motion capture
multi-person motion capture
0.812024
MultiPhys: Multi-Person Physics-Aware 3D Motion Estimation · CVPR 2024
Wearable and physiological sensing › wearable display
smart glasses
0.312025
VQA-Driven Event Maps for Assistive Navigation for People with Low Vision in Urban Environments · ICRA 2025
Computer vision › 3D vision
scene flow estimation
0.112011
Scene flow estimation by growing correspondence seeds · CVPR 2011
Computer vision › 3D vision
camera pose estimation
0.112010
Simultaneous pose, correspondence and non-rigid shape · CVPR 2010
Computer vision › 3D vision › 3d reconstruction
non-rigid reconstruction
0.112010
Simultaneous pose, correspondence and non-rigid shape · CVPR 2010
Computer vision › 3D vision › correspondence estimation
optical flow and stereo matching
0.012011
Scene flow estimation by growing correspondence seeds · CVPR 2011

Methods — techniques the papers use, named apart from their topics

visual question answering · 1.7semantic element extraction · 1.7physics simulation · 0.8autoregressive modeling · 0.8seed growing · 0.1correspondence propagation · 0.1kalman filter · 0.1gaussian mixture · 0.1correspondence estimation · 0.1
YearPublicationVenuePosition
2025 VQA-Driven Event Maps for Assistive Navigation for People with Low Vision in Urban Environments
abstract
We introduce a novel framework for assistive urban navigation for individuals with low vision. Utilizing a smart glasses platform developed by Biel Glasses, which provide a continuous stream of stereo images and GPS fixes, we generate an Event Map based on key semantic elements extracted by carefully prompted visual question-answering (VQA) models. For individuals with blurry or reduced fields of vision (low vision), traversing city streets poses a variety of challenges; they may struggle to perceive construction work, potholes, crowded sidewalks, and other ambiguous obstacles obstructing their paths. Some tasks, such as distinguishing traffic light signals, are nigh impossible without assistance from a companion or city infrastructure aimed towards accessibility. Although the majority of these problems may be solved with individually tailored traditional computer vision algorithms, developing and running a suite of these algorithms is challenging and resource demanding. Therefore, our proposed solution capitalizes on a single underlying implementation that need only be extended by adding queries. We validate our approach using a custom dataset of over 1,300 annotated images from various locations around Barcelona, reporting performance across different urban navigation tasks. We demonstrate the performance of the end to end system on a run of data collected by the Biel Glasses platform.
Joseph Morales, Bruk Gebregziabher, Alex Cabañeros, Jordi Sanchez-Riera
ICRA4
2024 InstantAvatar: Efficient 3D Head Reconstruction via Surface Rendering
abstract
Recent advances in full-head reconstruction have been obtained by optimizing a neural field through differentiable surface or volume rendering to represent a single scene. While these techniques achieve an unprecedented accuracy, they take several minutes, or even hours, due to the expensive optimization process required. In this work, we introduce InstantAvatar, a method that recovers full-head avatars from few images (down to just one) in a few seconds on commodity hardware. In order to speed up the reconstruction process, we propose a system that combines, for the first time, a voxel-grid neural field representation with a surface renderer. Notably, a naive combination of these two techniques leads to unstable optimizations that do not converge to valid solutions. In order to overcome this limitation, we present a novel statistical model that learns a prior distribution over 3D head signed distance functions using a voxel-grid based architecture. The use of this prior model, in combination with other design choices, results into a system that achieves 3D head reconstructions with comparable accuracy as the state-of-the-art with a 100 $\times$ speed-up.
Antonio Canela, Pol Caselles, Ibrar Malik, Eduard Ramon, Jaime García 0001, Jordi Sanchez-Riera, Gil Triginer, Francesc Moreno-Noguer
3DV6
2024 MultiPhys: Multi-Person Physics-Aware 3D Motion Estimation
abstract
We introduce MultiPhys, a method designed for recovering multi-person motion from monocular videos. Our focus lies in capturing coherent spatial placement between pairs of individuals across varying degrees of engagement. MultiPhys, being physically aware, exhibits robustness to jittering and occlusions, and effectively eliminates penetration issues between the two individuals. We devise a pipeline in which the motion estimated by a kinematic-based method is fed into a physics simulator in an autoregressive manner. We introduce distinct components that enable our model to harness the simulator's properties without compromising the accuracy of the kinematic estimates. This results in final motion estimates that are both kinematically coherent and physically compliant. Extensive evaluations on three challenging datasets characterized by substantial inter-person interaction show that our method significantly reduces errors associated with penetration and foot skating, while performing competitively with the state-of-the-art on motion accuracy and smoothness. Results and code can be found in our project page.
Nicolas Ugrinovic, Boxiao Pan, Georgios Pavlakos, Despoina Paschalidou, Bokui Shen, Jordi Sanchez-Riera, Francesc Moreno-Noguer, Leonidas J. Guibas
CVPR6
2021 PhysXNet: A Customizable Approach for Learning Cloth Dynamics on Dressed People
abstract
We introduce PhysXNet, a learning-based approach to predict the dynamics of deformable clothes given 3D skeleton motion sequences of humans wearing these clothes. The proposed model is adaptable to a large variety of garments and changing topologies, without need of being retrained. Such simulations are typically carried out by physics engines that require manual human expertise and are subject to computationally intensive computations. PhysXNet, by contrast, is a fully differentiable deep network that at inference is able to estimate the geometry of dense cloth meshes in a matter of milliseconds, and thus, can be readily deployed as a layer of a larger deep learning architecture. This efficiency is achieved thanks to the specific parameterization of the clothes we consider, based on 3D UV maps encoding spatial garment displacements. The problem is then formulated as a mapping between the human kinematics space (represented also by 3D UV maps of the undressed body mesh) into the clothes displacement UV maps, which we learn using a conditional GAN with a discriminator that enforces feasible deformations. We train simultaneously our model for three garment templates, tops, bottoms and dresses for which we simulate deformations under 50 different human actions. Nevertheless, the UV map representation we consider allows encapsulating many different cloth topologies, and at test we can simulate garments even if we did not specifically train for them. A thorough evaluation demonstrates that PhysXNet delivers cloth deformations very close to those computed with the physical engine, opening the door to be effectively integrated within deep learning pipelines.
Jordi Sanchez-Riera, Albert Pumarola, Francesc Moreno-Noguer
3DV1
2018 Robust RGB-D Hand Tracking Using Deep Learning Priors
abstract
With the irruption of inexpensive depth sensor devices, hand gesture tracking has become a topic of great interest. Two main problems to face respect other tracking algorithms are the high complexity of the hand structure, which translate in a very large amount of possible gestures, and the rapidness of the movements we are able to make when moving the hand or just the fingers. Recent approaches try to fit a 3D hand model to the observed RGB-D data by an optimization function that minimizes the error between the model and the data. However, these algorithms are very dependent on the initialization point, which are impractical to run in a natural environment. To solve these kinds of problems, it is common to use an offline data set with prelearned gestures that will serve as a first rough estimate. In concrete, we present an algorithm that uses an articulated ICP minimization function that is initialized by the parameters obtained from a data set of hand gestures trained through a deep learning framework. This setup has two strong points. First, deep learning provides a very fast and accurate estimate of performed hand gestures. Second, the articulated ICP algorithm allows capturing the possible variability of a gesture performed by different persons or slightly different gestures. Our proposed algorithm is evaluated and validated in several ways. Independent evaluations for the deep learning framework and articulated ICP are performed. Moreover, different real sequences are recorded to validate our approach and, finally, quantitative and qualitative comparisons are conducted with state-of-the-art algorithms.
Jordi Sanchez-Riera, Kathiravan Srinivasan, Kai-Lung Hua, Wen-Huang Cheng, M. Anwar Hossain 0001, Mohammed F. Alhamid
IEEE Trans. Circuits Syst. Video Technol.1
2017 i-Stylist: Finding the Right Dress Through Your Social Networks
Jordi Sanchez-Riera, Jun-Ming Lin, Kai-Lung Hua, Wen-Huang Cheng, Arvin Wen Tsui
MMM (1)1
2016 A comparative study of data fusion for RGB-D based visual recognition
Jordi Sanchez-Riera, Kai-Lung Hua, Yuan-Sheng Hsiao, Tekoing Lim, Shintami Chusnul Hidayati, Wen-Huang Cheng
Pattern Recognit. Lett.1
2014 LaRED: a large RGB-D extensible hand gesture dataset
abstract
We present the LaRED, a Large RGB-D Extensible hand gesture Dataset, recorded with an Intel's newly-developed short range depth camera. This dataset is unique and differs from the existing ones in several aspects. Firstly, the large volume of data recorded: 243, 000 tuples where each tuple is composed of a color image, a depth image, and a mask of the hand region. Secondly, the number of different classes provided: a total of 81 classes (27 gestures in 3 different rotations). Thirdly, the extensibility of dataset: the software used to record and inspect the dataset is also available, giving the possibility for future users to increase the number of data as well as the number of gestures. Finally, in this paper, some experiments are presented to characterize the dataset and establish a baseline as the start point to develop more complex recognition algorithms. The LaRED dataset is publicly available at: http://mclab.citi.sinica.edu.tw/dataset/lared/lared.html.
Yuan-Sheng Hsiao, Jordi Sanchez-Riera, Tekoing Lim, Kai-Lung Hua, Wen-Huang Cheng
MMSys2
2013 Benchmarking methods for audio-visual recognition using tiny training sets
abstract
The problem of choosing a classifier for audio-visual command recognition is addressed. Because such commands are culture- and user-dependant, methods need to learn new commands from a few examples. We benchmark three state-of-the-art discriminative classifiers based on bag of words and SVM. The comparison is made on monocular and monaural recordings of a publicly available dataset. We seek for the best trade off between speed, robustness and size of the training set. In the light of over 150,000 experiments, we conclude that this is a promising direction of work towards a flexible methodology that must be easily adaptable to a large variety of users.
Xavier Alameda-Pineda, Jordi Sanchez-Riera, Radu Horaud
ICASSP2
2012 Audio-visual robot command recognition: D-META'12 grand challenge
abstract
This paper addresses the problem of audio-visual command recognition in the framework of the D-META Grand Challenge1. Temporal and non-temporal learning models are trained on visual and auditory descriptors. In order to set a proper baseline, the methods are tested on the "Robot Gestures" scenario of the publicly available RAVEL data set, following the leave-one-out cross-validation strategy. The classification-level audio-visual fusion strategy allows for compensating the errors of the unimodal (audio or vision) classifiers. The obtained results (an average audio-visual recognition rate of almost 80%) encourage us to investigate on how to further develop and improve the methodology described in this paper.
Jordi Sanchez-Riera, Xavier Alameda-Pineda, Radu Horaud
ICMI1
2012 Robust spatiotemporal stereo for dynamic scenes
Jordi Sanchez-Riera, Jan Cech, Radu Horaud
ICPR1
2011 Scene flow estimation by growing correspondence seeds
abstract
A simple seed growing algorithm for estimating scene flow in a stereo setup is presented. Two calibrated and synchronized cameras observe a scene and output a sequence of image pairs. The algorithm simultaneously computes a disparity map between the image pairs and optical flow maps between consecutive images. This, together with calibration data, is an equivalent representation of the 3D scene flow, i.e. a 3D velocity vector is associated with each reconstructed point. The proposed method starts from correspondence seeds and propagates these correspondences to their neighborhood. It is accurate for complex scenes with large motions and produces temporally-coherent stereo disparity and optical flow results. The algorithm is fast due to inherent search space reduction. An explicit comparison with recent methods of spatiotemporal stereo and variational optical and scene flow is provided.
Jan Cech, Jordi Sanchez-Riera, Radu Horaud
CVPR2
2010 Simultaneous pose, correspondence and non-rigid shape
abstract
Recent works have shown that 3D shape of non-rigid surfaces can be accurately retrieved from a single image given a set of 3D-to-2D correspondences between that image and another one for which the shape is known. However, existing approaches assume that such correspondences can be readily established, which is not necessarily true when large deformations produce significant appearance changes between the input and the reference images. Furthermore, it is either assumed that the pose of the camera is known, or the estimated solution is pose-ambiguous. In this paper we relax all these assumptions and, given a set of 3D and 2D unmatched points, we present an approach to simultaneously solve their correspondences, compute the camera pose and retrieve the shape of the surface in the input image. This is achieved by introducing weak priors on the pose and shape that we model as Gaussian Mixtures. By combining them into a Kalman filter we can progressively reduce the number of 2D candidates that can be potentially matched to each 3D point, while pose and shape are refined. This lets us to perform a complete and efficient exploration of the solution space and retain the best solution.
Jordi Sanchez-Riera, Jonas Östlund, Pascal Fua, Francesc Moreno-Noguer
CVPR1
2010 Feature distribution modelling techniques for 3D face verification
Chris McCool, Jordi Sanchez-Riera, Sébastien Marcel
Pattern Recognit. Lett.2