Sergio Orts

dblp:92/9685 · also Sergio Orts-Escolano · DBLP profile ↗
← Back
61ranked-venue papers
8as first author
15since 2021 · last 2024
0000-0001-6817-6326ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 MagicMirror: Fast and High-Quality Avatar Generation with a Constrained Search Space
Armand Comas Massague, Di Qiu, Menglei Chai, Marcel C. Bühler, Amit Raj, Ruiqi Gao, Qiangeng Xu, Mark Matthews, Paulo F. U. Gotardo, Sergio Orts, Thabo Beeler
ECCV (66)10
2024 Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot Captures
abstract
Volumetric modeling and neural radiance field representations have revolutionized 3D face capture and photorealistic novel view synthesis. However, these methods often require hundreds of multi-view input images and are thus inapplicable to cases with less than a handful of inputs. We present a novel volumetric prior on human faces that allows for high-fidelity expressive face modeling from as few as three input views captured in the wild. Our key insight is that an implicit prior trained on synthetic data alone can generalize to extremely challenging real-world identities and expressions and render novel views with fine idiosyncratic details like wrinkles and eyelashes. We leverage a 3D Morphable Face Model to synthesize a large training set, rendering each identity with different expressions, hair, clothing, and other assets. We then train a conditional Neural Radiance Field prior on this synthetic dataset and, at inference time, fine-tune the model on a very sparse set of real images of a single subject. On average, the fine-tuning requires only three inputs to cross the synthetic-to-real domain gap. The resulting personalized 3D model reconstructs strong idiosyncratic facial expressions and outperforms the state-of-the-art in high-quality novel view synthesis of faces from sparse inputs in terms of perceptual and photo-metric quality.
Marcel C. Bühler, Gengyan Li 0001, Erroll Wood, Leonhard Helminger, Xu Chen 0025, Tanmay Shah, Daoye Wang, Stephan J. Garbin, Sergio Orts, Otmar Hilliges, Dmitry Lagun, Jérémy Riviere, Paulo F. U. Gotardo, Thabo Beeler, Abhimitra Meka, Kripasindhu Sarkar
SIGGRAPH Asia9
2024 GANtlitz: Ultra High Resolution Generative Model for Multi-Modal Face Textures
abstract
Abstract High‐resolution texture maps are essential to render photoreal digital humans for visual effects or to generate data for machine learning. The acquisition of high resolution assets at scale is cumbersome, it involves enrolling a large number of human subjects, using expensive multi‐view camera setups, and significant manual artistic effort to align the textures. To alleviate these problems, we introduce GANtlitz (A play on the german noun Antlitz, meaning face), a generative model that can synthesize multi‐modal ultra‐high‐resolution face appearance maps for novel identities. Our method solves three distinct challenges: 1) unavailability of a very large data corpus generally required for training generative models, 2) memory and computational limitations of training a GAN at ultra‐high resolutions, and 3) consistency of appearance features such as skin color, pores and wrinkles in high‐resolution textures across different modalities. We introduce dual‐style blocks, an extension to the style blocks of the StyleGAN2 architecture, which improve multi‐modal synthesis. Our patch‐based architecture is trained only on image patches obtained from a small set of face textures (<100) and yet allows us to generate seamless appearance maps of novel identities at 6k × 4k resolution. Extensive qualitative and quantitative evaluations and baseline comparisons show the efficacy of our proposed system. (see https://www.acm.org/publications/class-2012 )
Aurel Gruber, Edo Collins, Abhimitra Meka, Franziska Mueller 0001, Kripasindhu Sarkar, Sergio Orts, Luca Prasso, Jay Busch, Markus Gross 0001, Thabo Beeler
Comput. Graph. Forum6
2023 Rapsai: Accelerating Machine Learning Prototyping of Multimedia Applications through Visual Programming
abstract
In recent years, there has been a proliferation of multimedia applications that leverage machine learning (ML) for interactive experiences. Prototyping ML-based applications is, however, still challenging, given complex workflows that are not ideal for design and experimentation. To better understand these challenges, we conducted a formative study with seven ML practitioners to gather insights about common ML evaluation workflows.
Ruofei Du, Na Li 0034, Michelle Carney, Scott Miles, Maria Kleiner, Xiuxiu Yuan, Yinda Zhang 0001, Anuva Kulkarni, Xingyu Liu 0002, Ahmed Sabie, Sergio Orts, Abhishek Kar, Ram Iyengar, Adarsh Kowdle, Alex Olwal
CHI12
2023 Learning Personalized High Quality Volumetric Head Avatars from Monocular RGB Videos
abstract
We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracking of a 3DMM with a neural radiance field to achieve fine-grained control and photorealism. To reduce over-smoothing and improve out-of-model expressions synthesis, we propose to predict local features anchored on the 3DMM geometry. These learnt features are driven by 3DMM deformation and interpolated in 3D space to yield the volumetric radiance at a designated query point. We further show that using a Convolutional Neural Network in the UV space is critical in incorporating spatial context and producing representative local features. Extensive experiments show that we are able to reconstruct high-quality avatars, with more accurate expression-dependent details, good generalization to out-of-training expressions, and quantitatively superior renderings compared to other state-of-the-art approaches.
Ziqian Bai, Feitong Tan, Zeng Huang, Kripasindhu Sarkar, Danhang Tang, Di Qiu, Abhimitra Meka, Ruofei Du, Mingsong Dou, Sergio Orts, Rohit Pandey, Ping Tan 0002, Thabo Beeler, Sean Ryan Fanello, Yinda Zhang 0001
CVPR10
2023 Controllable Light Diffusion for Portraits
abstract
We introduce light diffusion, a novel method to improve lighting in portraits, softening harsh shadows and specular highlights while preserving overall scene illumi-nation. Inspired by professional photographers' diffusers and scrims, our method softens lighting given only a single portrait photo. Previous portrait relighting approaches focus on changing the entire lighting environment, removing shadows (ignoring strong specular highlights), or removing shading entirely. In contrast, we propose a learning based method that allows us to control the amount of light diffusion and apply it on in-the-wild portraits. Additionally, we design a method to synthetically generate plausible external shadows with sub-surface scattering effects while conforming to the shape of the subject's face. Finally, we show how our approach can increase the robustness of higher level vision applications, such as albedo estimation, geometry estimation and semantic segmentation.
David Futschik, Kelvin Ritland, James Vecore, Sean Ryan Fanello, Sergio Orts, Brian Curless, Daniel Sýkora, Rohit Pandey
CVPR5
2023 Preface: A Data-driven Volumetric Prior for Few-shot Ultra High-resolution Face Synthesis
abstract
NeRFs have enabled highly realistic synthesis of human faces including complex appearance and reflectance effects of hair and skin. These methods typically require a large number of multi-view input images, making the process hardware intensive and cumbersome, limiting applicability to unconstrained settings. We propose a novel volumetric human face prior that enables the synthesis of ultra high-resolution novel views of subjects that are not part of the prior’s training distribution. This prior model consists of an identity-conditioned NeRF, trained on a dataset of low-resolution multi-view images of diverse humans with known camera calibration. A simple sparse landmark-based 3D alignment of the training dataset allows our model to learn a smooth latent space of geometry and appearance despite a limited number of training identities. A high-quality volumetric representation of a novel subject can be obtained by model fitting to 2 or 3 camera views of arbitrary resolution. Importantly, our method requires as few as two views of casually captured images as input at inference time.
Marcel C. Bühler, Kripasindhu Sarkar, Tanmay Shah, Gengyan Li 0001, Daoye Wang, Leonhard Helminger, Sergio Orts, Dmitry Lagun, Otmar Hilliges, Thabo Beeler, Abhimitra Meka
ICCV7
2023 LitNeRF: Intrinsic Radiance Decomposition for High-Quality View Synthesis and Relighting of Faces
abstract
High-fidelity, photorealistic 3D capture of a human face is a long-standing problem in computer graphics – the complex material of skin, intricate geometry of hair, and fine scale textural details make it challenging. Traditional techniques rely on very large and expensive capture rigs to reconstruct explicit mesh geometry and appearance maps, and are limited by the accuracy of hand-crafted reflectance models. More recent volumetric methods (e.g., NeRFs) have enabled view-synthesis and sometimes relighting by learning an implicit representation of the density and reflectance basis, but suffer from artifacts and blurriness due to the inherent ambiguities in volumetric modeling. These problems are further exacerbated when capturing with few cameras and light sources. We present a novel technique for high-quality capture of a human face for 3D view synthesis and relighting using a sparse, compact capture rig consisting of 15 cameras and 15 lights. Our method combines a neural volumetric representation with traditional mesh reconstruction from multiview stereo. The proxy geometry allows us to anchor the 3D density field to prevent artifacts and guide the disentanglement of intrinsic radiance components of the face appearance such as diffuse and specular reflectance, and incident radiance (shadowing) fields. Our hybrid representation significantly improves the state-of-the-art quality for arbitrarily dense renders of a face from desired camera viewpoint as well as environmental, directional, and near-field lighting.
Kripasindhu Sarkar, Marcel C. Bühler, Gengyan Li 0001, Daoye Wang, Delio Vicini, Jérémy Riviere, Yinda Zhang 0001, Sergio Orts, Paulo F. U. Gotardo, Thabo Beeler, Abhimitra Meka
SIGGRAPH Asia8
2022 Synthetic contact maps to predict grasp regions on objects
abstract
From early ages, we learn to naturally grasp objects from our trial-and-error interaction with the environment. We end up excelling at manipulating objects, such that grabbing a pan by its handle is all but unconscious. However, performing the same task from a machine's perspective is particularly challenging, as it requires an in-depth understanding of object manipulation. We could approximate such an understanding by capturing the hand-object contact resulting from human grasps, i.e., contact maps. This has already been done using cumbersome capture systems, and not on a large scale. In this work, we simplify and accelerate such a process by generating contact maps from our interaction with household objects in photorealistic virtual reality environments. To the best of our knowledge, we are the first to generate contact maps at scale from our interaction in a virtual scenario. We train an image-to-image translation method to predict grasp regions on objects, demonstrating the usefulness of our generated contact maps. We provide all the necessary tools, code, and dataset to foster applications where hand-object understanding is necessary.
Pablo Martinez-Gonzalez, David Mulero-Pérez, Sergiu Ovidiu-Oprea, Manuel Benavent-Lledó, Sergio Orts, José García Rodríguez 0001
IJCNN5
2022 A Review on Deep Learning Techniques for Video Prediction
abstract
The ability to predict, anticipate and reason about future outcomes is a key component of intelligent decision-making systems. In light of the success of deep learning in computer vision, deep-learning-based video prediction emerged as a promising research direction. Defined as a self-supervised learning task, video prediction represents a suitable framework for representation learning, as it demonstrated potential capabilities for extracting meaningful representations of the underlying patterns in natural videos. Motivated by the increasing interest in this task, we provide a review on the deep learning methods for prediction in video sequences. We first define the video prediction fundamentals, as well as mandatory background concepts and the most used datasets. Next, we carefully analyze existing video prediction models organized according to a proposed taxonomy, highlighting their contributions and their significance in the field. The summary of the datasets and methods is accompanied with experimental results that facilitate the assessment of the state of the art on a quantitative basis. The paper is summarized by drawing some general conclusions, identifying open research challenges and by pointing out future research directions.
Sergiu Ovidiu-Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts, José García Rodríguez 0001, Antonis A. Argyros
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 NVS-MonoDepth: Improving Monocular Depth Prediction with Novel View Synthesis
abstract
Building upon the recent progress in novel view synthesis, we propose its application to improve monocular depth estimation. In particular, we propose a novel training method split in three main steps. First, the prediction results of a monocular depth network are warped to an additional view point. Second, we apply an additional image synthesis network, which corrects and improves the quality of the warped RGB image. The output of this network is required to look as similar as possible to the ground-truth view by minimizing the pixel-wise RGB reconstruction error. Third, we reapply the same monocular depth estimation onto the synthesized second view point and ensure that the depth predictions are consistent with the associated ground truth depth. Experimental results prove that our method achieves state-of-the-art or comparable performance on the KITTI and NYU-Depth-v2 datasets with a lightweight and simple vanilla U-Net architecture.
Zuria Bauer, Zuoyue Li, Sergio Orts, Miguel Cazorla, Marc Pollefeys, Martin R. Oswald
3DV3
2021 Graph Convolutional Neural Networks-based 3D Hand Pose Estimation over Point Clouds
abstract
In recent years we can find a multitude of approaches that aim to return the 3D pose of the hands. Most of them try to estimate the pose from RGB images or even include some geometrical information via depth maps. Furthermore, some proposals have shown promising results using point clouds as input data. However, the sparse nature of this type of data is often one of its drawbacks. To tackle this sparsity, different strategies have been brought to the table such as voxelizing or sorting the input data to impose a structure to the input domain. In this paper, we address this problem by means of a graph structure. This process implies that we should accommodate the point cloud onto a graph representation that connects its points. We connect each point to its neighborhood, a method that has been successfully used in similar proposals and whose clustering effect enables us to emulate an effect similar to kernels in image convolutions. The proposed architecture uses both graph and 2D convolutions. The first one aims to extract local features and build a feature map, from which the 2D convolutions will extract a second level of features used to estimate the pose. This proposal shows initial results to return a 3D pose of the hand from depth maps, which are projected on point clouds and redefined as graphs. Although the results diverge from other more established methods in the state of the art, it presents a proof of concept by which to address this problem without losing spatial information.
John Alejandro Castro-Vargas, Pablo Martinez-Gonzalez, Sergiu Ovidiu-Oprea, Alberto Garcia-Garcia, Sergio Orts, José García Rodríguez 0001
IJCNN5
2021 UnrealROX+: An Improved Tool for Acquiring Synthetic Data from Virtual 3D Environments
abstract
Synthetic data generation has become essential in last years for feeding data-driven algorithms, which surpassed traditional techniques performance in almost every computer vision problem. Gathering and labelling the amount of data needed for these data-hungry models in the real world may become unfeasible and error-prone, while synthetic data give us the possibility of generating huge amounts of data with pixel-perfect annotations. However, most synthetic datasets lack from enough realism in their rendered images. In that context UnrealROX generation tool was presented in 2019, allowing to generate highly realistic data, at high resolutions and framerates, with an efficient pipeline based on Unreal Engine, a cutting-edge videogame engine. UnrealROX enabled robotic vision researchers to generate realistic and visually plausible data with full ground truth for a wide variety of problems such as class and instance semantic segmentation, object detection, depth estimation, visual grasping, and navigation. Nevertheless, its workflow was very tied to generate image sequences from a robotic on-board camera, making hard to generate data for other purposes. In this work, we present UnrealROX+, an improved version of UnrealROX where its decoupled and easy-to-use data acquisition system allows to quickly design and generate data in a much more flexible and customizable way. Moreover, it is packaged as an Unreal plug-in, which makes it more comfortable to use with already existing Unreal projects, and it also includes new features such as generating albedo or a Python API for interacting with the virtual environment from Deep Learning frameworks.
Pablo Martinez-Gonzalez, Sergiu Ovidiu-Oprea, John Alejandro Castro-Vargas, Alberto Garcia-Garcia, Sergio Orts, José García Rodríguez 0001, Markus Vincze
IJCNN5
2021 H-GAN: the power of GANs in your Hands
abstract
We present HandGAN (H-GAN), a cycle-consistent adversarial learning approach implementing multi-scale perceptual discriminators. It is designed to translate synthetic images of hands to the real domain. Synthetic hands provide complete ground-truth annotations, yet they are not representative of the target distribution of real-world data. We strive to provide the perfect blend of a realistic hand appearance with synthetic annotations. Relying on image-to-image translation, we improve the appearance of synthetic hands to approximate the statistical distribution underlying a collection of real images of hands. H-GAN tackles not only the cross-domain tone mapping but also structural differences in localized areas such as shading discontinuities. Results are evaluated on a qualitative and quantitative basis improving previous works. Furthermore, we relied on the hand classification task to claim our generated hands are statistically similar to the real domain of hands.
Sergiu Ovidiu-Oprea, Giorgos Karvounas, Pablo Martinez-Gonzalez, Nikolaos Kyriazis, Sergio Orts, Iasonas Oikonomidis, Alberto Garcia-Garcia, Aggeliki Tsoli, José García Rodríguez 0001, Antonis A. Argyros
IJCNN5
2021 Total relighting: learning to relight portraits for background replacement
abstract
We propose a novel system for portrait relighting and background replacement, which maintains high-frequency boundary details and accurately synthesizes the subject's appearance as lit by novel illumination, thereby producing realistic composite images for any desired scene. Our technique includes foreground estimation via alpha matting, relighting, and compositing. We demonstrate that each of these stages can be tackled in a sequential pipeline without the use of priors (e.g. known background or known illumination) and with no specialized acquisition techniques, using only a single RGB portrait image and a novel, target HDR lighting environment as inputs. We train our model using relit portraits of subjects captured in a light stage computational illumination system, which records multiple lighting conditions, high quality geometry, and accurate alpha mattes. To perform realistic relighting for compositing, we introduce a novel per-pixel lighting representation in a deep learning framework, which explicitly models the diffuse and the specular components of appearance, producing relit portraits with convincingly rendered non-Lambertian effects like specular highlights. Multiple experiments and comparisons show the effectiveness of the proposed approach when applied to in-the-wild images.
Rohit Pandey, Sergio Orts, Chloe LeGendre, Christian Häne, Sofien Bouaziz, Christoph Rhemann, Paul E. Debevec, Sean Ryan Fanello
ACM Trans. Graph.2
2020 Du2Net: Learning Depth Estimation from Dual-Cameras and Dual-Pixels
Yinda Zhang 0001, Neal Wadhwa, Sergio Orts, Christian Häne, Sean Ryan Fanello, Rahul Garg 0002
ECCV (1)3
2020 Enhancing perception for the visually impaired with deep learning techniques and low-cost wearable sensors
abstract
As estimated by the World Health Organization, there are millions of people who lives with some form of vision impairment . As a consequence, some of them present mobility problems in outdoor environments . With the aim of helping them, we propose in this work a system which is capable of delivering the position of potential obstacles in outdoor scenarios. Our approach is based on non-intrusive wearable devices and focuses also on being low-cost. First, a depth map of the scene is estimated from a color image, which provides 3D information of the environment. Then, an urban object detector is in charge of detecting the semantics of the objects in the scene. Finally, the three-dimensional and semantic data is summarized in a simpler representation of the potential obstacles the users have in front of them. This information is transmitted to the user through spoken or haptic feedback. Our system is able to run at about 3.8 fps and achieved a 87.99% mean accuracy in obstacle presence detection. Finally, we deployed our system in a pilot test which involved an actual person with vision impairment, who validated the effectiveness of our proposal for improving its navigation capabilities in outdoors.
Zuria Bauer, Alejandro Dominguez, Edmanuel Cruz, Francisco Gomez-Donoso, Sergio Orts, Miguel Cazorla
Pattern Recognit. Lett.5
2020 COMBAHO: A deep learning system for integrating brain injury patients in society
abstract
In the last years, the care of dependent people, either by disease, accident, disability, or age, is one of the current priority research topics in developed countries. Moreover, such care is intended to be at patients home, in order to minimize the cost of therapies. Patients rehabilitation will be fulfilled when their integration in society is achieved, either in the family or in a work environment. To address this challenge, we propose the development and evaluation of an assistant for people with acquired brain injury or dependents. This assistant is twofold: in the patient’s home is based on the design and use of an intelligent environment with abilities to monitor and active learning, combined with an autonomous social robot for interactive assistance and stimulation. On the other hand, it is complemented with an outdoor assistant, to help patients under disorientation or complex situations. This involves the integration of several existing technologies and provides solutions to a variety of technological challenges. Deep leaning-based techniques are proposed as core technology to solve these problems.
José García Rodríguez 0001, Francisco Gomez-Donoso, Sergiu Ovidiu-Oprea, Alberto Garcia-Garcia, Miguel Cazorla, Sergio Orts, Zuria Bauer, John Alejandro Castro-Vargas, Félix Escalona, David Ivorra-Piqueres, Pablo Martinez-Gonzalez, Eugenio Aguirre, Miguel García-Silvente, Marcelo García-Pérez, José María Cañas, Francisco Martín 0001, Jonatan Gines Clavero, Francisco Rivas-Montero
Pattern Recognit. Lett.6
2020 Deep relightable textures: volumetric performance capture with neural rendering
abstract
The increasing demand for 3D content in augmented and virtual reality has motivated the development of volumetric performance capture systemsnsuch as the Light Stage. Recent advances are pushing free viewpoint relightable videos of dynamic human performances closer to photorealistic quality. However, despite significant efforts, these sophisticated systems are limited by reconstruction and rendering algorithms which do not fully model complex 3D structures and higher order light transport effects such as global illumination and sub-surface scattering. In this paper, we propose a system that combines traditional geometric pipelines with a neural rendering scheme to generate photorealistic renderings of dynamic performances under desired viewpoint and lighting. Our system leverages deep neural networks that model the classical rendering process to learn implicit features that represent the view-dependent appearance of the subject independent of the geometry layout, allowing for generalization to unseen subject poses and even novel subject identity. Detailed experiments and comparisons demonstrate the efficacy and versatility of our method to generate high-quality results, significantly outperforming the existing state-of-the-art solutions.
Abhimitra Meka, Rohit Pandey, Christian Häne, Sergio Orts, Peter Barnum, Philip L. Davidson, Daniel Erickson, Yinda Zhang 0001, Jonathan Taylor 0001, Sofien Bouaziz, Chloe LeGendre, Wan-Chun Ma, Ryan S. Overbeck, Thabo Beeler, Paul E. Debevec, Shahram Izadi, Christian Theobalt, Christoph Rhemann, Sean Ryan Fanello
ACM Trans. Graph.4
2019 TactileGCN: A Graph Convolutional Network for Predicting Grasp Stability with Tactile Sensors
abstract
Tactile sensors provide useful contact data during the interaction with an object which can be used to accurately learn to determine the stability of a grasp. Most of the works in the literature represented tactile readings as plain feature vectors or matrix-like tactile images, using them to train machine learning models. In this work, we explore an alternative way of exploiting tactile information to predict grasp stability by leveraging graph-like representations of tactile data, which preserve the actual spatial arrangement of the sensor's taxels and their locality. In experimentation, we trained a Graph Neural Network to binary classify grasps as stable or slippery ones. To train such network and prove its predictive capabilities for the problem at hand, we captured a novel dataset of ~ 5000 three-fingered grasps across 41 objects for training and 1000 grasps with 10 unknown objects for testing. Our experiments prove that this novel approach can be effectively used to predict grasp stability.
Alberto Garcia-Garcia, Brayan S. Zapata-Impata, Sergio Orts, Pablo Gil, José García Rodríguez 0001
IJCNN3
2019 A visually realistic grasping system for object manipulation and interaction in virtual reality environments
Sergiu Ovidiu-Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts, José García Rodríguez 0001
Comput. Graph.5
2019 Accurate and efficient 3D hand pose regression for robot hand teleoperation using a monocular RGB camera
Francisco Gomez-Donoso, Sergio Orts, Miguel Cazorla
Expert Syst. Appl.2
2019 Large-scale multiview 3D hand pose dataset
Francisco Gomez-Donoso, Sergio Orts, Miguel Cazorla
Image Vis. Comput.2
2019 Evaluation of different chrominance models in the detection and reconstruction of faces and hands using the growing neural gas network
abstract
Physical traits such as the shape of the hand and face can be used for human recognition and identification in video surveillance systems and in biometric authentication smart card systems, as well as in personal health care. However, the accuracy of such systems suffers from illumination changes, unpredictability, and variability in appearance (e.g. occluded faces or hands, cluttered backgrounds, etc.). This work evaluates different statistical and chrominance models in different environments with increasingly cluttered backgrounds where changes in lighting are common and with no occlusions applied, in order to get a reliable neural network reconstruction of faces and hands, without taking into account the structural and temporal kinematics of the hands. First a statistical model is used for skin colour segmentation to roughly locate hands and faces. Then a neural network is used to reconstruct in 3D the hands and faces. For the filtering and the reconstruction we have used the growing neural gas algorithm which can preserve the topology of an object without restarting the learning process. Experiments conducted on our own database but also on four benchmark databases (Stirling's, Alicante, Essex, and Stegmann's) and on deaf individuals from normal 2D videos are freely available on the BSL signbank dataset. Results demonstrate the validity of our system to solve problems of face and hand segmentation and reconstruction under different environmental conditions.
Anastassia Angelopoulou, José García Rodríguez 0001, Sergio Orts, Epaminondas Kapetanios, Xing Liang, Bencie Woll, Alexandra Psarrou
Pattern Anal. Appl.3
2019 The relightables: volumetric performance capture of humans with realistic relighting
abstract
We present "The Relightables", a volumetric capture system for photorealistic and high quality relightable full-body performance capture. While significant progress has been made on volumetric capture systems, focusing on 3D geometric reconstruction with high resolution textures, much less work has been done to recover photometric properties needed for relighting. Results from such systems lack high-frequency details and the subject's shading is prebaked into the texture. In contrast, a large body of work has addressed relightable acquisition for image-based approaches, which photograph the subject under a set of basis lighting conditions and recombine the images to show the subject as they would appear in a target lighting environment. However, to date, these approaches have not been adapted for use in the context of a high-resolution volumetric capture system. Our method combines this ability to realistically relight humans for arbitrary environments, with the benefits of free-viewpoint volumetric capture and new levels of geometric accuracy for dynamic performances. Our subjects are recorded inside a custom geodesic sphere outfitted with 331 custom color LED lights, an array of high-resolution cameras, and a set of custom high-resolution depth sensors. Our system innovates in multiple areas: First, we designed a novel active depth sensor to capture 12.4 MP depth maps, which we describe in detail. Second, we show how to design a hybrid geometric and machine learning reconstruction pipeline to process the high resolution input and output a volumetric video. Third, we generate temporally consistent reflectance maps for dynamic performers by leveraging the information contained in two alternating color gradient illumination images acquired at 60Hz. Multiple experiments, comparisons, and applications show that The Relightables significantly improves upon the level of realism in placing volumetrically captured human performances into arbitrary CG scenes.
Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts, Rohit Pandey, Jason Dourgarian, Danhang Tang, Anastasia Tkach, Adarsh Kowdle, Emily Cooper, Mingsong Dou, Sean Ryan Fanello, Graham Fyffe, Christoph Rhemann, Jonathan Taylor 0001, Paul E. Debevec, Shahram Izadi
ACM Trans. Graph.8
2018 A New Dataset and Performance Evaluation of a Region-based CNN for Urban Object Detection
abstract
In the last years, we have seen a large growth in the number of applications which use deep learning-based object detectors. Autonomous Driving Assistance Systems (ADAS) is one of the areas where it has more impact. In this work, we present a novel study that evaluates a state-of-the-art technique for urban object localization. In particular, we investigate the performance of the Faster R-CNN method to detect and localize urban objects in a variety of outdoor urban videos involving pedestrians, cars, bicycles and other objects moving in the scene. We propose a new dataset that is used for benchmarking the accuracy of a real-time object detector (Faster R-CNN). Part of the data was collected using an HD camera mounted in a vehicle. Besides, some of the data is weakly annotated so it can be used for testing weakly-supervised learning techniques. We have carried out extensive experiments demonstrating the effectiveness of the baseline approach, which achieved a 74.2% accuracy on the proposed dataset. Moreover, we have evaluated a baseline approach for traffic sign recognition achieving an accuracy of 98.1%. A ResNet-based architecture was trained and used for this purpose as a second stage of our object detector. The full dataset is available for download at http://www.rovit.ua.es/dataset/traffic/.
Alex Dominguez-Sanchez, Sergio Orts, José García Rodríguez 0001, Miguel Cazorla
IJCNN2
2018 The RobotriX: An Extremely Photorealistic and Very-Large-Scale Indoor Dataset of Sequences with Robot Trajectories and Interactions
abstract
Enter the RobotriX, an extremely photorealistic indoor dataset designed to enable the application of deep learning techniques to a wide variety of robotic vision problems. The RobotriX consists of hyperrealistic indoor scenes which are explored by robot agents which also interact with objects in a visually realistic manner in that simulated world. Photorealistic scenes and robots are rendered by Unreal Engine into a virtual reality headset which captures gaze so that a human operator can move the robot and use controllers for the robotic hands; scene information is dumped on a per-frame basis so that it can be reproduced offline using UnrealCV to generate raw data and ground truth labels. By taking this approach, we were able to generate a dataset of 38 semantic classes across 512 sequences totaling 8M stills recorded at +60 frames per second with full HD resolution. For each frame, RGB-D and 3D information is provided with full annotations in both spaces. Thanks to the high quality and quantity of both raw information and annotations, the RobotriX will serve as a new milestone for investigating 2D and 3D robotic vision tasks with large-scale data-driven techniques.
Alberto Garcia-Garcia, Pablo Martinez-Gonzalez, Sergiu Ovidiu-Oprea, John Alejandro Castro-Vargas, Sergio Orts, José García Rodríguez 0001, Alvaro Jover-Alvarez
IROS5
2018 A long short-term memory based Schaeffer gesture recognition system
abstract
Abstract In this work, a Schaeffer language recognition system is proposed in order to help autistic children overcome communicative disorders. Using Schaeffer language as a speech and language therapy, improves children communication skills and at the same time the understanding of language productions. Nevertheless, the teaching process of children in performing gestures properly is not straightforward. For this purpose, this system will teach children with autism disorder the correct way to communicate using gestures in combination with speech reproduction. The main purpose is to accelerate the learning process and increase children interest by using a technological approach. Several recurrent neural network‐based approaches have been tested, such as vanilla recurrent neural networks, long short‐term memory networks,and gated recurrent unit‐based models. In order to select the most suitable model, an extensive comparison has been conducted reporting a 93.13% classification success rate over a subset of 25 Schaeffer gestures by using an long short‐term memory‐based approach. Our dataset consists on pose‐based features such as angles and euclidean distances extracted from the raw skeletal data provided by a Kinect v2 sensor.
Sergiu Ovidiu-Oprea, Alberto Garcia-Garcia, Sergio Orts, Victor Villena-Martinez, John Alejandro Castro-Vargas
Expert Syst. J. Knowl. Eng.3
2018 Fast 2D/3D object representation with growing neural gas
abstract
This work presents the design of a real-time system to model visual objects with the use of self-organising networks. The architecture of the system addresses multiple computer vision tasks such as image segmentation, optimal parameter estimation and object representation. We first develop a framework for building non-rigid shapes using the growth mechanism of the self-organising maps, and then we define an optimal number of nodes without overfitting or underfitting the network based on the knowledge obtained from information-theoretic considerations. We present experimental results for hands and faces, and we quantitatively evaluate the matching capabilities of the proposed method with the topographic product. The proposed method is easily extensible to 3D objects, as it offers similar features for efficient mesh reconstruction.
Anastassia Angelopoulou, José García Rodríguez 0001, Sergio Orts, Gaurav Gupta 0001, Alexandra Psarrou
Neural Comput. Appl.3
2018 Bioinspired point cloud representation: 3D object tracking
Sergio Orts, José García Rodríguez 0001, Miguel Cazorla, Vicente Morell, Jorge Azorín López, Marcelo Saval-Calvo, Alberto Garcia-Garcia, Victor Villena-Martinez
Neural Comput. Appl.1
2017 LonchaNet: A sliced-based CNN architecture for real-time 3D object recognition
abstract
In the last few years, Convolutional Neural Networks (CNNs) had become the default paradigm to address classification problems, specially, but not only, in image recognition. This is mainly due to the high success rate that they provide. Despite there currently exist approaches that apply deep learning to the 3D recognition problem, they are either too slow for online uses or too error prone. To fill this gap, we propose LonchaNet, a deep learning architecture for point clouds classification. Our system successfully achieves a high accuracy yet providing a low computation cost. A dense set of experiments were carried out in order to validate our system in the frame of the ModelNet - a large-scale 3D CAD models dataset - challenge. Our proposal achieves a success rate of 94.37% in the ModelNet-10 classification task, the second place in the leaderboard as of today (November, 2016).
Francisco Gomez-Donoso, Alberto Garcia-Garcia, José García Rodríguez 0001, Sergio Orts, Miguel Cazorla
IJCNN4
2017 A recurrent neural network based Schaeffer gesture recognition system
abstract
Schaeffer language is considered an effective method to help autistic children overcome communicative disorders. Speech and language therapy results in an improvement in communication skills and understanding of language productions. In this work, a Schaeffer language recognition system is presented with the purpose of teaching children with autism disorder the correct way to communicate using gestures in combination with speech reproduction. The purpose is to accelerate the learning process and increase children interest using a technological approach. A Long Short-Term Memory (LSTM) model has been implemented for this purpose reporting a 93.13% classification success rate over a subset of 25 Schaeffer gestures. A comparison with vanilla RNNs and GRU-based models has been also carried out. Pose-based features such as angles and euclidean distances have been extracted from our gesture dataset by processing raw skeletal data from a Kinect v2 sensor.
Sergiu Ovidiu-Oprea, Alberto Garcia-Garcia, José García Rodríguez 0001, Sergio Orts, Miguel Cazorla
IJCNN4
2017 A study of the effect of noise and occlusion on the accuracy of convolutional neural networks applied to 3D object recognition
Alberto Garcia-Garcia, José García Rodríguez 0001, Sergio Orts, Sergiu Ovidiu-Oprea, Francisco Gomez-Donoso, Miguel Cazorla
Comput. Vis. Image Underst.3
2017 Interactive light source position estimation for augmented reality with an RGB-D camera
abstract
Abstract The first hybrid CPU‐GPU based method for estimating a point light source position in a scene recorded by an RGB‐D camera is presented. The image and depth information from the Kinect is enough to estimate a light position in a scene, which allows for the rendering of synthetic objects into a scene that appears realistic enough for augmented reality purposes. This method does not require a light probe or other physical device. To make this method suitable for augmented reality, we developed a hybrid implementation that performs light estimation in under 1second. This is sufficient for most augmented reality scenarios because both the position of the light source and the position of the Kinect are typically fixed. The method is able to estimate the angle of the light source with an average error of 20°. By rendering synthetic objects into the recorded scene, we illustrate that this accuracy is good enough for the rendered objects to look realistic. Copyright © 2015 John Wiley & Sons, Ltd.
Bas Boom, Sergio Orts, Xi Ning, Steven McDonagh 0001, Peter Sandilands, Robert B. Fisher
Comput. Animat. Virtual Worlds2
2017 Multi-sensor 3D object dataset for object recognition with full pose estimation
Alberto Garcia-Garcia, Sergio Orts, Sergiu Ovidiu-Oprea, José García Rodríguez 0001, Jorge Azorín López, Marcelo Saval-Calvo, Miguel Cazorla
Neural Comput. Appl.2
2017 Evaluation of sampling method effects in 3D non-rigid registration
Marcelo Saval-Calvo, Jorge Azorín López, Andrés Fuster Guilló, José García Rodríguez 0001, Sergio Orts, Alberto Garcia-Garcia
Neural Comput. Appl.5
2017 Object recognition in noisy RGB-D data using GNG
José Carlos Rangel, Vicente Morell, Miguel Cazorla, Sergio Orts, José García Rodríguez 0001
Pattern Anal. Appl.4
2017 A robotic platform for customized and interactive rehabilitation of persons with disabilities
Francisco Gomez-Donoso, Sergio Orts, Alberto Garcia-Garcia, José García Rodríguez 0001, John Alejandro Castro-Vargas, Sergiu Ovidiu-Oprea, Miguel Cazorla
Pattern Recognit. Lett.2
2017 Pedestrian Movement Direction Recognition Using Convolutional Neural Networks
abstract
Pedestrian movement direction recognition is an important factor in autonomous driver assistance and security surveillance systems. Pedestrians are the most crucial and fragile moving objects in streets, roads, and events, where thousands of people may gather on a regular basis. People flow analysis on zebra crossings and in shopping centers or events such as demonstrations are a key element to improve safety and to enable autonomous cars to drive in real life environments. This paper focuses on deep learning techniques such as convolutional neural networks (CNN) to achieve a reliable detection of pedestrians moving in a particular direction. We propose a CNN-based technique that leverages current pedestrian detection techniques (histograms of oriented gradients-linSVM) to generate a sum of subtracted frames (flow estimation around the detected pedestrian), which are used as an input for the proposed modified versions of various state-of-the-art CNN networks, such as AlexNet, GoogleNet, and ResNet. Moreover, we have also created a new data set for this purpose, and analyzed the importance of training in a known data set for the neural networks to achieve reliable results.
Alex Dominguez-Sanchez, Miguel Cazorla, Sergio Orts
IEEE Trans. Intell. Transp. Syst.3
2016 A Closed-Form Bayesian Fusion Equation Using Occupancy Probabilities
abstract
We present a new mathematical framework for multi-view surface reconstruction from a set of calibrated color and depth images. We estimate the occupancy probability of points in space along sight rays, and combine these estimates using a normalized product derived from Bayes' rule. The advantage of this approach is that the free space constraint is a natural consequence of the formulation, and not a separate logical operation. We present a single closed form implicit expression for the reconstructed surface in terms of the image data and camera projections, making analytic properties such as surface normals not only easy to compute, but exact. This expression can be efficiently evaluated on the GPU, making it ideal for high performance real-time applications, such as live human body capture for immersive telepresence.
Charles T. Loop, Qin Cai, Sergio Orts, Philip A. Chou
3DV3
2016 HyperDepth: Learning Depth from Structured Light without Matching
abstract
Structured light sensors are popular due to their robustness to untextured scenes and multipath. These systems triangulate depth by solving a correspondence problem between each camera and projector pixel. This is often framed as a local stereo matching task, correlating patches of pixels in the observed and reference image. However, this is computationally intensive, leading to reduced depth accuracy and framerate. We contribute an algorithm for solving this correspondence problem efficiently, without compromising depth accuracy. For the first time, this problem is cast as a classification-regression task, which we solve extremely efficiently using an ensemble of cascaded random forests. Our algorithm scales in number of disparities, and each pixel can be processed independently, and in parallel. No matching or even access to the corresponding reference pattern is required at runtime, and regressed labels are directly mapped to depth. Our GPU-based algorithm runs at a 1KHz for 1.3MP input/output images, with disparity error of 0.1 subpixels. We show a prototype high framerate depth camera running at 375Hz, useful for solving tracking-related problems. We demonstrate our algorithmic performance, creating high resolution real-time depth maps that surpass the quality of current state of the art depth technologies, highlighting quantization-free results with reduced holes, edge fattening and other stereo-based depth artifacts.
Sean Ryan Fanello, Christoph Rhemann, Vladimir Tankovich, Adarsh Kowdle, Sergio Orts, David Kim 0002, Shahram Izadi
CVPR5
2016 PointNet: A 3D Convolutional Neural Network for real-time object class recognition
abstract
During the last few years, Convolutional Neural Networks are slowly but surely becoming the default method solve many computer vision related problems. This is mainly due to the continuous success that they have achieved when applied to certain tasks such as image, speech, or object recognition. Despite all the efforts, object class recognition methods based on deep learning techniques still have room for improvement. Most of the current approaches do not fully exploit 3D information, which has been proven to effectively improve the performance of other traditional object recognition methods. In this work, we propose PointNet, a new approach inspired by VoxNet and 3D ShapeNets, as an improvement over the existing methods by using density occupancy grids representations for the input data, and integrating them into a supervised Convolutional Neural Network architecture. An extensive experimentation was carried out, using ModelNet - a large-scale 3D CAD models dataset - to train and test the system, to prove that our approach is on par with state-of-the-art methods in terms of accuracy while being able to perform recognition under real-time constraints.
Alberto Garcia-Garcia, Francisco Gomez-Donoso, José García Rodríguez 0001, Sergio Orts, Miguel Cazorla, Jorge Azorín López
IJCNN4
2016 Holoportation: Virtual 3D Teleportation in Real-time
abstract
We present an end-to-end system for augmented and virtual reality telepresence, called Holoportation. Our system demonstrates high-quality, real-time 3D reconstructions of an entire space, including people, furniture and objects, using a set of new depth cameras. These 3D models can also be transmitted in real-time to remote users. This allows users wearing virtual or augmented reality displays to see, hear and interact with remote participants in 3D, almost as if they were present in the same physical space. From an audio-visual perspective, communicating and interacting with remote users edges closer to face-to-face communication. This paper describes the Holoportation technical system in full, its key interactive capabilities, the application scenarios it enables, and an initial qualitative study of using this new communication medium.
Sergio Orts, Christoph Rhemann, Sean Ryan Fanello, Wayne Chang, Adarsh Kowdle, Yury Degtyarev, David Kim 0002, Philip Davidson, Sameh Khamis, Mingsong Dou, Vladimir Tankovich, Charles T. Loop, Qin Cai, Philip A. Chou, Sarah Mennicken, Julien P. C. Valentin, Vivek Pradeep, Shenlong Wang, Sing Bing Kang, Pushmeet Kohli, Yuliya Lutchyn, Cem Keskin, Shahram Izadi
UIST1
2016 3D Surface Reconstruction of Noisy Point Clouds Using Growing Neural Gas: 3D Object/Scene Reconstruction
Sergio Orts, José García Rodríguez 0001, Vicente Morell, Miguel Cazorla, Jose Antonio Serra-Perez, Alberto Garcia-Garcia
Neural Process. Lett.1
2016 Fusion4D: real-time performance capture of challenging scenes
abstract
We contribute a new pipeline for live multi-view performance capture, generating temporally coherent high-quality reconstructions in real-time. Our algorithm supports both incremental reconstruction, improving the surface estimation over time, as well as parameterizing the nonrigid scene motion. Our approach is highly robust to both large frame-to-frame motion and topology changes, allowing us to reconstruct extremely challenging scenes. We demonstrate advantages over related real-time techniques that either deform an online generated template or continually fuse depth data nonrigidly into a single reference model. Finally, we show geometric reconstruction results on par with offline methods which require orders of magnitude more processing time and many more RGBD cameras.
Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts, Christoph Rhemann, David Kim 0002, Jonathan Taylor 0001, Pushmeet Kohli, Vladimir Tankovich, Shahram Izadi
ACM Trans. Graph.7
2015 Self-Organizing Activity Description Map to represent and classify human behaviour
abstract
The automated understanding of people activities from video sequences is an open research topic in which the computer vision and pattern recognition areas have made big efforts in recent years. This paper proposes the Self Organizing Activity Description Map (SOADM). It is a novel neural network based on the self-organizing paradigm to classify high level of semantic understanding from video sequences. The neural network is able to deal with the big gap between human trajectories in a scene and the global behaviour associated to them. Specifically, using simple representations of people trajectories as input, the SOADM is able to both represent and classify human behaviours. Additionally, the map is able to preserve the topological information about the scene. Experiments have been carried out using the Shopping Centre dataset of the CAVIAR database taken into account the global behaviour of an individual. Results confirm the high accuracy of the proposal outperforming previous methods.
Jorge Azorín López, Marcelo Saval-Calvo, Andrés Fuster Guilló, José García Rodríguez 0001, Sergio Orts
IJCNN5
2015 Processing point cloud sequences with Growing Neural Gas
abstract
We consider the problem of processing point cloud sequences. In particular, we represent and track objects in dynamic scenes acquired using low-cost sensors such as the Kinect. A neural network based approach is proposed to represent and estimate 3D objects motion. This system addresses multiple computer vision tasks such as object segmentation, representation, motion analysis and tracking. The use of a neural network allows the unsupervised estimation of motion and the representation of objects in the scene. This proposal avoids the problem of finding corresponding features while tracking moving objects. A set of experiments are presented that demonstrate the validity of our method to track 3D objects. Favorable results are presented demonstrating the capabilities of the GNG algorithm for this task.
Sergio Orts, José García Rodríguez 0001, Vicente Morell, Miguel Cazorla, Marcelo Saval-Calvo, Jorge Azorín López
IJCNN1
2015 Using GNG on 3D Object Recognition in Noisy RGB-D data
abstract
The object recognition task on 3D scenes is a growing research field that faces some problems relative to the use of 3D point clouds. In this work, we focus on dealing with the noise in the clouds through the use of the Growing Neural Gas (GNG) network filtering algorithm. The GNG method is able to represent the input data with a desired amount of neurons while preserving the topology of the input space. The selected recognition pipeline works describing extracted keypoints of the clouds, grouping and comparing it to detect the presence of an object in the scene, through a hypothesis verification algorithm. Experiments show how the GNG method yields better recognitions results that others filtering algorithms when noise is present.
José Carlos Rangel, Vicente Morell, Miguel Cazorla, Sergio Orts, José García Rodríguez 0001
IJCNN4
2015 Non-rigid point set registration using color and data downsampling
abstract
Nowadays, non-rigid registration problem is an active research topic in computer vision. Various proposals exist which face the problem from different perspectives, but it is still a challenging problem. Currently, with the new low-cost RGB-D sensors, the use of both, color and 3D information, is getting more interest in many applications. In this paper, we present a non-rigid registration technique based on CPD, and including color information along with 3D data, to estimate the non-rigid transformation. As the input data size is critical in the processing time, a sampling technique is required. Five sampling techniques are evaluated: a bilinear sampling, a normal-based, a color-based, a combination of the normal and color-based samplings, and a Growing Neural Gas based approach. All of them have been evaluated with the already presented non-rigid registration methods. Results show the performance of each sampling method, obtaining better results for the registration process using color-based sampling techniques.
Marcelo Saval-Calvo, Sergio Orts, Jorge Azorín López, José García Rodríguez 0001, Andrés Fuster Guilló, Vicente Morell, Miguel Cazorla
IJCNN2
2015 3D reconstruction of medical images from slices automatically landmarked with growing neural models
Anastassia Angelopoulou, Alexandra Psarrou, José García Rodríguez 0001, Sergio Orts, Jorge Azorín López, Kenneth Revett
Neurocomputing4
2014 3D maps representation using GNG
abstract
Current RGB-D sensors provide a big amount of valuable information for mobile robotics tasks like 3D map reconstruction, but the storage and processing of the incremental data provided by the different sensors through time quickly becomes unmanageable. In this work, we focus on 3D maps representation and we propose the use of a Growing Neural Gas (GNG) network as a 3D representation model of the input data. GNG method is able to represent the input data with a desired amount of neurons while preserving the topology of the input space. Experiments show how GNG method yields better input space adaptation than other state-of-the-art 3D map representation methods.
Vicente Morell, Miguel Cazorla, Sergio Orts, José García Rodríguez 0001
IJCNN3
2014 3D colour object reconstruction based on Growing Neural Gas
abstract
With the advent of low-cost 3D sensors and 3D printers, surface reconstruction has become an important research topic in the last years. In this work, we propose an automatic method for 3D surface reconstruction from raw unorganized point clouds acquired using low-cost sensors. We have modified the Growing Neural Gas (GNG) network, which is a suitable model because of its flexibility, rapid adaptation and excellent quality of representation, to perform 3D surface reconstruction of different real-world objects. Some improvements have been made on the original algorithm considering colour information during the learning stage and creating complete triangular meshes instead of basic wire-frame representations. The proposed method is able to create 3D faces online, whereas existing 3D reconstruction methods based on Self-Organizing Maps (SOMs) required post-processing steps to close gaps and holes produced during the 3D reconstruction process. Performed experiments validated how the proposed method improves existing techniques removing post-processing steps and including colour information in the final triangular mesh.
Sergio Orts, José García Rodríguez 0001, Vicente Morell, Miguel Cazorla, Juan Manuel García Chamizo
IJCNN1
2014 Geometric 3D point cloud compression
Vicente Morell, Sergio Orts, Miguel Cazorla, José García Rodríguez 0001
Pattern Recognit. Lett.2
2013 Point Light Source Estimation based on Scenes Recorded by a RGB-D camera
abstract
Estimation of the point light source position in the scene enhances the experience for augmented reality. The image and depth information from the RGB-D camera allows estimation of the point light source position in a scene, where our approach does not need any probe objects or other measuring devices. The approach uses the Lambertian reflectance model, where the RGB-D camera provides the image and the surface model and the remaining unknowns are the albedo and light parameters (light intensity and direction). In order to determine the light parameters, we assume that segments with a similar colour have the same albedo, which allows us to find the point light source that explains the illumination in the scene. The performance of this method is evaluated on multiple scenes, where a single light bulb is used to illuminate the scene. In this case, the average error in the angle between the true light position vector and our estimate is around 10 degrees. This allows realistic rendering of synthetic objects into the recorded scene, which is used to improve the experience of augmented reality.
Bas Boom, Sergio Orts, Xi Ning, Steven McDonagh 0001, Peter Sandilands, Robert B. Fisher
BMVC2
2013 Point cloud data filtering and downsampling using growing neural gas
abstract
3D sensors provide valuable information for mobile robotic tasks like scene classification or object recognition, but these sensors often produce noisy data that makes impossible applying classical keypoint detection and feature extraction techniques. Therefore, noise removal and downsampling have become essential steps in 3D data processing. In this work, we propose the use of a 3D filtering and downsampling technique based on a Growing Neural Gas (GNG) network. GNG method is able to deal with outliers presents in the input data. These features allows to represent 3D spaces, obtaining an induced Delaunay Triangulation of the input space. Experiments show how GNG method yields better input space adaptation to noisy data than other filtering and downsampling methods like Voxel Grid. It is also demonstrated how the state-of-the-art keypoint detectors improve their performance using filtered data with GNG network. Descriptors extracted on improved keypoints perform better matching in robotics applications as 3D scene registration.
Sergio Orts, Vicente Morell, José García Rodríguez 0001, Miguel Cazorla
IJCNN1
2013 Improving drug discovery using a neural networks based parallel scoring function
abstract
Virtual Screening (VS) methods can considerably aid clinical research, predicting how ligands interact with drug targets. Most VS methods suppose a unique binding site for the target, but it has been demonstrated that diverse ligands interact with unrelated parts of the target and many VS methods do not take into account this relevant fact. This problem is circumvented by a novel VS methodology named BINDSURF that scans the whole protein surface to find new hotspots, where ligands might potentially interact with, and which is implemented in massively parallel Graphics Processing Units, allowing fast processing of large ligand databases. BINDSURF can thus be used in drug discovery, drug design, drug repurposing and therefore helps considerably in clinical research. However, the accuracy of most VS methods is constrained by limitations in the scoring function that describes biomolecular interactions, and even nowadays these uncertainties are not completely understood. In order to solve this problem, we propose a novel approach where neural networks are trained with databases of known active (drugs) and inactive compounds, and later used to improve VS predictions.
Horacio Emilio Pérez Sánchez, Ginés D. Guerrero, José M. García 0001, Jorge Peña-García, José M. Cecilia, Gaspar Cano, Sergio Orts, José García Rodríguez 0001
IJCNN7
2013 3D gesture recognition with growing neural gas
abstract
We propose the design of a real-time system to recognize and interpret hand gestures. The acquisition devices are low cost 3D sensors. 3D hand pose segmentation, characterization and tracking will be implemented using the growing neural gas (GNG) structure. The capacity of the system to obtain information with a high degree of freedom allows the encoding of many gestures and a very accurate motion capture. The use of hand pose models combined with motion information provided with GNG permits to deal with the problem of the hand motion representation. A natural interface applied to a virtual mirror writing system and a module to estimate hand pose have been designed to demonstrate the validity of the system.
Jose Antonio Serra-Perez, José García Rodríguez 0001, Sergio Orts, Juan Manuel García Chamizo, Javier Montoyo-Bojo, Anastassia Angelopoulou, Alexandra Psarrou, Markos Mentzelopoulos, Andrew Lewis 0004
IJCNN3
2012 Multi-GPU based camera network system keeps privacy using growing neural gas
abstract
In this work we present a multi-camera surveillance system based on the use of self-organizing neural networks to represent events in video. The objectives include: identifying and tracking persons or objects in the scene or the interpretation of user gestures for interaction with services, devices and systems implemented in the digital home. Additionally, the system process several tasks in parallel using GPUs (Graphic Processor Units). Addressing multiple vision tasks of various levels such as segmentation, representation or characterization, analysis and monitoring of the movement to allow the construction of a robust representation of their environment and interpret the elements of the scene.
Sergio Orts, José García Rodríguez 0001, Vicente Morell, Jorge Azorín López, Juan Manuel García Chamizo
IJCNN1
2012 GPGPU implementation of growing neural gas: Application to 3D scene reconstruction
Sergio Orts, José García Rodríguez 0001, Diego Viejo, Miguel Cazorla, Vicente Morell
J. Parallel Distributed Comput.1
2012 Autonomous Growing Neural Gas for applications with time constraint: Optimal parameter estimation
José García Rodríguez 0001, Anastassia Angelopoulou, Juan Manuel García Chamizo, Alexandra Psarrou, Sergio Orts, Vicente Morell
Neural Networks5
2011 Fast Autonomous Growing Neural Gas
abstract
This paper aims to address the ability of self-organizing neural network models to manage real-time applications. Specifically, we introduce fAGNG (fast Autonomous Growing Neural Gas), a modified learning algorithm for the incremental model Growing Neural Gas (GNG) network. The Growing Neural Gas network with its attributes of growth, flexibility, rapid adaptation, and excellent quality of representation of the input space makes it a suitable model for real time applications. However, under time constraints GNG fails to produce the optimal topological map for any input data set. In contrast to existing algorithms the proposed fAGNG algorithm introduces multiple neurons per iteration. The number of neurons inserted and input data generated is controlled autonomous and dynamically based on a priory learnt model. Comparative experiments using topological preservation measures are carried out to demonstrate the effectiveness of the new algorithm to represent linear and non-linear input spaces under time restrictions.
José García Rodríguez 0001, Anastassia Angelopoulou, Juan Manuel García Chamizo, Alexandra Psarrou, Sergio Orts, Vicente Morell
IJCNN5