Javier Hidalgo-Carrió

dblp:255/5486 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Robot navigation and mapping · 46% 3D vision · 42% Segmentation and scene understanding · 12%
Computer graphics and multimedia
1 paper
Computational photography and imaging · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
0.612022
Event-aided Direct Sparse Odometry · CVPR 2022
Robotics › Robot navigation and mapping › visual odometry
direct sparse odometry
0.612022
Event-aided Direct Sparse Odometry · CVPR 2022
Robotics › Robot navigation and mapping › visual odometry
event-based visual odometry
0.612022
Event-aided Direct Sparse Odometry · CVPR 2022
Computer vision › 3D vision › 3d reconstruction
photometric bundle adjustment
0.612022
Event-aided Direct Sparse Odometry · CVPR 2022
Robotics › Robot navigation and mapping
visual odometry
0.612022
Event-aided Direct Sparse Odometry · CVPR 2022
Computer vision › 3D vision
event-based vision
0.412020
Video to Events: Recycling Video Datasets for Event Cameras · CVPR 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.412020
Video to Events: Recycling Video Datasets for Event Cameras · CVPR 2020
Computational photography and imaging › event-based vision
event camera simulation
0.412020
Video to Events: Recycling Video Datasets for Event Cameras · CVPR 2020

Methods — techniques the papers use, named apart from their topics

video-to-event conversion · 0.9probabilistic brightness increment estimation · 0.6event generation model · 0.6
YearPublicationVenuePosition
2023 Autonomous Power Line Inspection with Drones via Perception-Aware MPC
abstract
Drones have the potential to revolutionize power line inspection by increasing productivity, reducing inspection time, improving data quality, and eliminating the risks for human operators. Current state-of-the-art systems for power line inspection have two shortcomings: (i) control is decoupled from perception and needs accurate information about the location of the power lines and masts; (ii) obstacle avoidance is decoupled from the power line tracking, which results in poor tracking in the vicinity of the power masts, and, consequently, in decreased data quality for visual inspection. In this work, we propose a model predictive controller (MPC) that overcomes these limitations by tightly coupling perception and action. Our controller generates commands that maximize the visibility of the power lines while, at the same time, safely avoiding the power masts. For power line detection, we propose a lightweight learning-based detector that is trained only on synthetic data and is able to transfer zero-shot to real-world power line images. We validate our system in simulation and real-world experiments on a mock-up power line infrastructure. We release our code and datasets to the public.
Jiaxu Xing, Giovanni Cioffi, Javier Hidalgo-Carrió, Davide Scaramuzza 0001
IROS3
2022 Event-aided Direct Sparse Odometry
abstract
We introduce EDS, a direct monocular visual odometry using events and frames. Our algorithm leverages the event generation model to track the camera motion in the blind time between frames. The method formulates a direct probabilistic approach of observed brightness increments. Per-pixel brightness increments are predicted using a sparse number of selected 3D points and are compared to the events via the brightness increment error to estimate camera motion. The method recovers a semi-dense 3D map using photometric bundle adjustment. EDS is the first method to perform 6-DOF VO using events and frames with a direct approach. By design it overcomes the problem of changing appearance in indirect methods. Our results outperform all previous event-based odometry solutions. We also show that, for a target error performance, EDS can work at lower frame rates than state-of-the-art frame-based VO solutions. This opens the door to low-power motion-tracking applications where frames are sparingly triggered “on demand” and our method tracks the motion in between. We release code and datasets to the public.
Javier Hidalgo-Carrió, Guillermo Gallego 0002, Davide Scaramuzza 0001
CVPR1
2021 Powerline Tracking with Event Cameras
abstract
Autonomous inspection of powerlines with quadrotors is challenging. Flights require persistent perception to keep a close look at the lines. We propose a method that uses event cameras to robustly track powerlines. Event cameras are inherently robust to motion blur, have low latency, and high dynamic range. Such properties are advantageous for autonomous inspection of powerlines with drones, where fast motions and challenging illumination conditions are ordinary. Our method identifies lines in the stream of events by detecting planes in the spatio-temporal signal, and tracks them through time. The implementation runs onboard and is capable of detecting multiple distinct lines in real time with rates of up to 320 thousand events per second. The performance is evaluated in real-world flights along a powerline. The tracker is able to persistently track the powerlines, with a mean lifetime of the line 10× longer than existing approaches.
Alex Dietsche, Giovanni Cioffi, Javier Hidalgo-Carrió, Davide Scaramuzza 0001
IROS3
2020 Learning Monocular Dense Depth from Events
abstract
Event cameras are novel sensors that output brightness changes in the form of a stream of asynchronous ”events” instead of intensity frames. Compared to conventional image sensors, they offer significant advantages: high temporal resolution, high dynamic range, no motion blur, and much lower bandwidth. Recently, learning-based approaches have been applied to event-based data, thus unlocking their potential and making significant progress in a variety of tasks, such as monocular depth prediction. Most existing approaches use standard feed-forward architectures to generate network predictions, which do not leverage the temporal consistency presents in the event stream. We propose a recurrent architecture to solve this task and show significant improvement over standard feed-forward methods. In particular, our method generates dense depth predictions using a monocular setup, which has not been shown previously. We pretrain our model using a new dataset containing events and depth maps recorded in the CARLA simulator. We test our method on the Multi Vehicle Stereo Event Camera Dataset (MVSEC). Quantitative experiments show up to 50% improvement in average depth error with respect to previous event-based methods. Code and dataset are available at: http://rpg.ifi.uzh.ch/e2depth.
Javier Hidalgo-Carrió, Daniel Gehrig, Davide Scaramuzza 0001
3DV1
2020 Video to Events: Recycling Video Datasets for Event Cameras
abstract
Event cameras are novel sensors that output brightness changes in the form of a stream of asynchronous "events" instead of intensity frames. They offer significant advantages with respect to conventional cameras: high dynamic range (HDR), high temporal resolution, and no motion blur. Recently, novel learning approaches operating on event data have achieved impressive results. Yet, these methods require a large amount of event data for training, which is hardly available due the novelty of event sensors in computer vision research. In this paper, we present a method that addresses these needs by converting any existing video dataset recorded with conventional cameras to synthetic event data. This unlocks the use of a virtually unlimited number of existing video datasets for training networks designed for real event data. We evaluate our method on two relevant vision tasks, i.e., object recognition and semantic segmentation, and show that models trained on synthetic events have several benefits: (i) they generalize well to real event data, even in scenarios where standard-camera images are blurry or overexposed, by inheriting the outstanding properties of event cameras; (ii) they can be used for fine-tuning on real data to improve over state-of-the-art for both classification and semantic segmentation.
Daniel Gehrig, Mathias Gehrig, Javier Hidalgo-Carrió, Davide Scaramuzza 0001
CVPR3