Hernán Badino

dblp:46/6186 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Systems, architecture and hardware · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Face, body and person analysis · 66% 3D vision · 26% Robot navigation and mapping · 6%
Computer graphics and multimedia
5 papers
Virtual and augmented reality · 51% Computer animation and physical simulation · 39% Geometric modeling and processing · 10%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d human pose estimation
1.022023
SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2023
xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera · ICCV 2019
Computer vision › Face, body and person analysis › human pose estimation › 3d pose estimation
egocentric pose estimation
1.022023
SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2023
xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera · ICCV 2019
Computer vision › Face, body and person analysis
human pose estimation
1.022023
SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2023
xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera · ICCV 2019
Virtual and augmented reality
avatar
0.612022
Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual Reality · CVPR 2022
Computer animation and physical simulation › facial animation
facial expression transfer
0.612022
Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual Reality · CVPR 2022
Computer vision › Face, body and person analysis
face tracking
0.412019
VR facial animation via multiview image translation · ACM Trans. Graph. 2019
Computer animation and physical simulation
facial animation
0.412019
VR facial animation via multiview image translation · ACM Trans. Graph. 2019
Virtual and augmented reality › virtual reality
social virtual reality
0.412019
VR facial animation via multiview image translation · ACM Trans. Graph. 2019
Virtual and augmented reality › immersive display
head-mounted display
0.322023
SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera · IEEE Trans. Pattern Anal. Mach. Intell. 2023
xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera · ICCV 2019
Computer vision › Face, body and person analysis
facial expression analysis
0.212022
Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual Reality · CVPR 2022
Robotics › Robot navigation and mapping
localization
0.112012
Real-time topometric localization · ICRA 2012
Geometric modeling and processing
point cloud processing
0.112011
Fast and accurate computation of surface normals from range images · ICRA 2011
Geometric modeling and processing › point cloud processing
range image processing
0.112011
Fast and accurate computation of surface normals from range images · ICRA 2011
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.112019
VR facial animation via multiview image translation · ACM Trans. Graph. 2019
Robotics › Robot navigation and mapping › localization › global localization
kidnapped robot problem
0.012012
Real-time topometric localization · ICRA 2012
Robotics › Robot navigation and mapping › terrain perception
terrain estimation
0.012011
Fast and accurate computation of surface normals from range images · ICRA 2011

Methods — techniques the papers use, named apart from their topics

synthetic data generation · 2.1encoder-decoder architecture · 2.1multi-branch decoder · 1.3unsupervised disentanglement · 1.1multi-identity architecture · 1.1data augmentation · 1.1self-supervised learning · 0.8multi-view geometry · 0.8end-to-end tracking · 0.83d features · 0.1spherical range image derivatives · 0.1
YearPublicationVenuePosition
2023 SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera
abstract
We present a new solution to egocentric 3D body pose estimation from monocular images captured from a downward looking fish-eye camera installed on the rim of a head mounted virtual reality device. This unusual viewpoint leads to images with unique visual appearance, characterized by severe self-occlusions and strong perspective distortions that result in a drastic difference in resolution between lower and upper body. We propose a new encoder-decoder architecture with a novel multi-branch decoder designed specifically to account for the varying uncertainty in 2D joint locations. Our quantitative evaluation, both on synthetic and real-world datasets, shows that our strategy leads to substantial improvements in accuracy over state of the art egocentric pose estimation approaches. To tackle the severe lack of labelled training data for egocentric 3D pose estimation we also introduced a large-scale photo-realistic synthetic dataset. xR-EgoPose offers 383K frames of high quality renderings of people with diverse skin tones, body shapes and clothing, in a variety of backgrounds and lighting conditions, performing a range of actions. Our experiments show that the high variability in our new synthetic training corpus leads to good generalization to real world footage and to state of the art results on real world datasets with ground truth. Moreover, an evaluation on the Human3.6M benchmark shows that the performance of our method is on par with top performing approaches on the more classic problem of 3D human pose from a third person viewpoint.
Denis Tomè, Thiemo Alldieck, Patrick Peluse, Gerard Pons-Moll, Lourdes Agapito, Hernán Badino, Fernando De la Torre
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual Reality
abstract
Social presence, the feeling of being there with a “real” person, will fuel the next generation of communication systems driven by digital humans in virtual reality (VR). The best 3D video-realistic VR avatars that minimize the uncanny effect rely on person-specific (PS) models. However, these PS models are time-consuming to build and are typically trained with limited data variability, which results in poor generalization and robustness. Major sources of variability that affects the accuracy of facial expression transfer algorithms include using different VR headsets (e.g., camera configuration, slop of the headset), facial appearance changes over time (e.g., beard, make-up), and environmental factors (e.g., lighting, backgrounds). This is a major drawback for the scalability of these models in VR. This paper makes progress in overcoming these limitations by proposing an end-to-end multi-identity architecture (MIA) trained with specialized augmentation strategies. MIA drives the shape component of the avatar from three cameras in the VR headset (two eyes, one mouth), in untrained subjects, using minimal personalized information (i.e., neutral 3D mesh shape). Similarly, if the PS texture decoder is available, MIA is able to drive the full avatar (shape + texture) robustly outperforming PS models in challenging scenarios. Our key contribution to improve robustness and generalization, is that our method implicitly decouples, in an unsupervised manner, the facial expression from nuisance factors (e.g., headset, environment, facial appearance). We demonstrate the superior performance and robustness of the proposed method versus state-of-the-art PS approaches in a variety of experiments.
Amin Jourabloo, Fernando De la Torre, Jason M. Saragih, Shih-En Wei, Stephen Lombardi, Te-Li Wang, Danielle Belko, Autumn Trimble, Hernán Badino
CVPR9
2019 xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera
abstract
We present a new solution to egocentric 3D body pose estimation from monocular images captured from a downward looking fish-eye camera installed on the rim of a head mounted virtual reality device. This unusual viewpoint, just 2 cm away from the user's face, leads to images with unique visual appearance, characterized by severe self-occlusions and strong perspective distortions that result in a drastic difference in resolution between lower and upper body. Our contribution is two-fold. Firstly, we propose a new encoder-decoder architecture with a novel dual branch decoder designed specifically to account for the varying uncertainty in the 2D joint locations. Our quantitative evaluation, both on synthetic and real-world datasets, shows that our strategy leads to substantial improvements in accuracy over state of the art egocentric pose estimation approaches. Our second contribution is a new large-scale photorealistic synthetic dataset - xR-EgoPose - offering 383K frames of high quality renderings ofpeople with a diversity of skin tones, body shapes, clothing, in a variety of backgrounds and lighting conditions, performing a range of actions. Our experiments show that the high variability in our new synthetic training corpus leads to good generalization to real world footage and to state of the art results on real world datasets with ground truth. Moreover, an evaluation on the Human3.6M benchmark shows that the performance of our method is on par with top performing approaches on the more classic problem of 3D human pose from a third person viewpoint.
Denis Tomè, Patrick Peluse, Lourdes Agapito, Hernán Badino
ICCV4
2019 VR facial animation via multiview image translation
abstract
A key promise of Virtual Reality (VR) is the possibility of remote social interaction that is more immersive than any prior telecommunication media. However, existing social VR experiences are mediated by inauthentic digital representations of the user (i.e., stylized avatars). These stylized representations have limited the adoption of social VR applications in precisely those cases where immersion is most necessary (e.g., professional interactions and intimate conversations). In this work, we present a bidirectional system that can animate avatar heads of both users' full likeness using consumer-friendly headset mounted cameras (HMC). There are two main challenges in doing this: unaccommodating camera views and the image-to-avatar domain gap. We address both challenges by leveraging constraints imposed by multiview geometry to establish precise image-to-avatar correspondence, which are then used to learn an end-to-end model for real-time tracking. We present designs for a training HMC, aimed at data-collection and model building, and a tracking HMC for use during interactions in VR. Correspondence between the avatar and the HMC-acquired images are automatically found through self-supervised multiview image translation, which does not require manual annotation or one-to-one correspondence between domains. We evaluate the system on a variety of users and demonstrate significant improvements over prior work.
Shih-En Wei, Jason M. Saragih, Tomas Simon, Adam W. Harley, Stephen Lombardi, Michal Perdoch, Alexander Hypes, Hernán Badino, Yaser Sheikh
ACM Trans. Graph.9
2014 Topometric localization on a road network
abstract
Current GPS-based devices have difficulty localizing in cases where the GPS signal is unavailable or insufficiently accurate. This paper presents an algorithm for localizing a vehicle on an arbitrary road network using vision, road curvature estimates, or a combination of both. The method uses an extension of topometric localization, which is a hybrid between topological and metric localization. The extension enables localization on a network of roads rather than just a single, non-branching route. The algorithm, which does not rely on GPS, is able to localize reliably in situations where GPS-based devices fail, including “urban canyons” in downtown areas and along ambiguous routes with parallel roads. We demonstrate the algorithm experimentally on several road networks in urban, suburban, and highway scenarios. We also evaluate the road curvature descriptor and show that it is effective when imagery is sparsely available.
Danfei Xu, Hernán Badino, Daniel F. Huber
IROS2
2014 Understanding how camera configuration and environmental conditions affect appearance-based localization
abstract
Localization is a central problem for intelligent vehicles. Visual localization can supplement or replace GPS-based localization approaches in situations where GPS is unavailable or inaccurate. Although visual localization has been demonstrated in a variety of algorithms and systems, the problem of how to best configure such a system remains largely an open question. Design choices, such as “where should the camera be placed?” and “how should it be oriented?” can have substantial effect on the cost and robustness of a fielded intelligent vehicle. This paper analyzes how different sensor configuration parameters and environmental conditions affect visual localization performance with the goal of understanding what causes certain configurations to perform better than others and providing general principles for configuring systems for visual localization. We ground the investigation using extensive field testing of a visual localization algorithm, and the data sets used for the analysis are made available for comparative evaluation.
Aayush Bansal, Hernán Badino, Daniel F. Huber
Intelligent Vehicles Symposium2
2012 Real-time topometric localization
abstract
Autonomous vehicles must be capable of localizing even in GPS denied situations. In this paper, we propose a real-time method to localize a vehicle along a route using visual imagery or range information. Our approach is an implementation of topometric localization, which combines the robustness of topological localization with the geometric accuracy of metric methods. We construct a map by navigating the route using a GPS-equipped vehicle and building a compact database of simple visual and 3D features. We then localize using a Bayesian filter to match sequences of visual or range measurements to the database. The algorithm is reliable across wide environmental changes, including lighting differences, seasonal variations, and occlusions, achieving an average localization accuracy of 1 m over an 8 km route. The method converges correctly even with wrong initial position estimates solving the kidnapped robot problem.
Hernán Badino, Daniel F. Huber, Takeo Kanade
ICRA1
2012 Improving sub-pixel accuracy for long range stereo
Stefan K. Gehrig, Hernán Badino, Uwe Franke
Comput. Vis. Image Underst.2
2011 Fast and accurate computation of surface normals from range images
abstract
The fast and accurate computation of surface normals from a point cloud is a critical step for many 3D robotics and automotive problems, including terrain estimation, mapping, navigation, object segmentation, and object recognition. To obtain the tangent plane to the surface at a point, the traditional approach applies total least squares to its small neighborhood. However, least squares becomes computationally very expensive when applied to the millions of measurements per second that current range sensors can generate. We reformulate the traditional least squares solution to allow the fast computation of surface normals, and propose a new approach that obtains the normals by calculating the derivatives of the surface from a spherical range image. Furthermore, we show that the traditional least squares problem is very sensitive to range noise and must be normalized to obtain accurate results. Experimental results with synthetic and real data demonstrate that our proposed method is not only more efficienThe fast and accurate computation of surface normals from a point cloud is a critical step for many 3D robotics and automotive problems, including terrain estimation, mapping, navigation, object segmentation, and object recognition. To obtain the tangent plane to the surface at a point, the traditional approach applies total least squares to its small neighborhood. However, least squares becomes computationally very expensive when applied to the millions of measurements per second that current range sensors can generate. We reformulate the traditional least squares solution to allow the fast computation of surface normals, and propose a new approach that obtains the normals by calculating the derivatives of the surface from a spherical range image. Furthermore, we show that the traditional least squares problem is very sensitive to range noise and must be normalized to obtain accurate results. Experimental results with synthetic and real data demonstrate that our proposed method is not only more efficient by up to two orders of magnitude, but provides better accuracy than the traditional least squares for practical neighborhood sizes.t by up to two orders of magnitude, but provides better accuracy than the traditional least squares for practical neighborhood sizes.
Hernán Badino, Daniel F. Huber, Yongwoon Park, Takeo Kanade
ICRA1
2011 Extrinsic calibration of a single line scanning lidar and a camera
abstract
In robotic hands design tendon driven systems have been considered for years. The main advantage is a small end effector inertia e.g. a light, small hand with high dynamics due to remote actuators. To protect the actuators from impact in unknown environments a compliant mechanism can be used. It absorbs energy during an impact or saves energy to enhance the joint dynamics. In this paper an antagonistic tendon mechanism is presented. It fits 38 times in the DLR Hand Arm System forearm and enables is adapted to the different finger joints and different tendon lengths. A magnetic sensor was developed for the force measurement of the tendons. Finally, the calibration and the robustness are demonstrated through a set of experiments.
Kiho Kwak, Daniel F. Huber, Hernán Badino, Takeo Kanade
IROS3
2011 Visual topometric localization
abstract
One of the fundamental requirements of an autonomous vehicle is the ability to determine its location on a map. Frequently, solutions to this localization problem rely on GPS information or use expensive three dimensional (3D) sensors. In this paper, we describe a method for long-term vehicle localization based on visual features alone. Our approach utilizes a combination of topological and metric mapping, which we call topometric localization, to encode the coarse topology of the route as well as detailed metric information required for accurate localization. A topometric map is created by driving the route once and recording a database of visual features. The vehicle then localizes by matching features to this database at runtime. Since individual feature matches are unreliable, we employ a discrete Bayes filter to estimate the most likely vehicle position using evidence from a sequence of images along the route. We illustrate the approach using an 8.8 km route through an urban and suburban environment. The method achieves an average localization error of 2.7 m over this route, with isolated worst case errors on the order of 10 m.
Hernán Badino, Daniel F. Huber, Takeo Kanade
Intelligent Vehicles Symposium1
2010 Latent Gaussian Mixture Regression for Human Pose Estimation
Leonid Sigal, Hernán Badino, Fernando De la Torre, Yong Liu 0027
ACCV (3)3
2009 B-Spline Modeling of Road Surfaces With an Application to Free-Space Estimation
abstract
We propose a general technique for modeling the visible road surface in front of a vehicle. The common assumption of a planar road surface is often violated in reality. A workaround proposed in the literature is the use of a piecewise linear or quadratic function to approximate the road surface. Our approach is based on representing the road surface as a general parametric B-spline curve. The surface parameters are tracked over time using a Kalman filter. The surface parameters are estimated from stereo measurements in the free space. To this end, we adopt a recently proposed road-obstacle segmentation algorithm to include disparity measurements and the B-spline road-surface representation. Experimental results in planar and undulating terrain verify the increase in free-space availability and accuracy using a flexible B-spline for road-surface modeling.
Andreas Wedel, Hernán Badino, Clemens Rabe, Heidi Loose, Uwe Franke, Daniel Cremers
IEEE Trans. Intell. Transp. Syst.2
2006 Accurate and Model-Free Pose Estimation of Small Objects for Crash Video Analysis
abstract
We propose a novel model-free pose estimation algorithm to estimate the relative pose of a rigid object. In most pose estimation algorithms, the object of interest covers a large portion of the image. We focus on pose estimation of small objects covering a field of view of less than 5° by 5° using stereo vision. With this new algorithm suitable for small objects, we investigate the effect of the object size on the pose accuracy. In addition, we introduce an object tracking technique that is insensitive to partial occlusion. The main application for this method is the analysis of crash video sequences. In this context, high-speed cameras with high resolution are used. However, the method easily extends to real-time robotics applications with VGA-size images, where relative pose estimation is needed, e.g. for manipulator control.
Stefan K. Gehrig, Hernán Badino, Pascal Paysan
BMVC2