Florian Golemo

dblp:08/8643 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
6since 2021 · last 2023
0000-0001-9238-7764ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 22% 3D vision · 19% Autonomous driving · 14%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
imitation learning
0.712023
Towards Learning to Imitate from a Single Video Demonstration · J. Mach. Learn. Res. 2023
Robotics › Robot manipulation › learning from demonstration
learning from video demonstration
0.712023
Towards Learning to Imitate from a Single Video Demonstration · J. Mach. Learn. Res. 2023
Machine learning › Reinforcement learning
reward learning
0.712023
Towards Learning to Imitate from a Single Video Demonstration · J. Mach. Learn. Res. 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.612022
Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction · ICLR 2022
Robotics › Autonomous driving › trajectory prediction
multi-agent trajectory prediction
0.612022
Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction · ICLR 2022
Computer vision › 3D vision › motion estimation
optical flow
0.612022
Kubric: A scalable dataset generator · CVPR 2022
Machine learning › Generative modeling
synthetic data generation
0.612022
Kubric: A scalable dataset generator · CVPR 2022
Robotics › Autonomous driving
trajectory prediction
0.612022
Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction · ICLR 2022
Machine learning › Reinforcement learning
differentiable simulation
0.512021
gradSim: Differentiable simulation for system identification and visuomotor control · ICLR 2021
Robotics › Motion planning and robot control
system identification
0.512021
gradSim: Differentiable simulation for system identification and visuomotor control · ICLR 2021
Computer vision › 3D vision
3d scene reconstruction
0.412020
Pix2Shape: Towards Unsupervised Learning of 3D Scenes from Images Using a View-Based Representation · Int. J. Comput. Vis. 2020
Machine learning › Representation and self-supervised learning
contrastive learning
0.412020
Unsupervised Learning of Dense Visual Representations · NeurIPS 2020
Computer vision › Segmentation and scene understanding
dense prediction
0.412020
Unsupervised Learning of Dense Visual Representations · NeurIPS 2020
Machine learning › Representation and self-supervised learning › contrastive learning › dense contrastive learning
pixel-level contrastive learning
0.412020
Unsupervised Learning of Dense Visual Representations · NeurIPS 2020
Computer vision › 3D vision
unsupervised 3d learning
0.412020
Pix2Shape: Towards Unsupervised Learning of 3D Scenes from Images Using a View-Based Representation · Int. J. Comput. Vis. 2020
Machine learning › Deep learning architectures and training › data-centric deep learning
training data generation
0.212022
Kubric: A scalable dataset generator · CVPR 2022
Robotics › Robot manipulation
visuomotor control
0.112021
gradSim: Differentiable simulation for system identification and visuomotor control · ICLR 2021
Computer vision › 3D vision › 3d shape representation › shape descriptor
view-based representation
0.112020
Pix2Shape: Towards Unsupervised Learning of 3D Scenes from Images Using a View-Based Representation · Int. J. Comput. Vis. 2020

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.1siamese recurrent neural network · 0.7reinforcement learning · 0.7transformer · 0.6set prediction · 0.6latent variable model · 0.6differentiable simulation · 0.5view-based representation · 0.4unsupervised learning · 0.4data augmentation · 0.4
YearPublicationVenuePosition
2023 Towards Learning to Imitate from a Single Video Demonstration
abstract
Agents that can learn to imitate behaviours observed in video -- without having direct access to internal state or action information of the observed agent -- are more suitable for learning in the natural world. However, formulating a reinforcement learning (RL) agent that facilitates this goal remains a significant challenge. We approach this challenge using contrastive training to learn a reward function by comparing an agent's behaviour with a single demonstration. We use a Siamese recurrent neural network architecture to learn rewards in space and time between motion clips while training an RL policy to minimize this distance. Through experimentation, we also find that the inclusion of multi-task data and additional image encoding losses improve the temporal consistency of the learned rewards and, as a result, significantly improve policy learning. We demonstrate our approach on simulated humanoid, dog, and raptor agents in 2D and quadruped and humanoid agents in 3D. We show that our method outperforms current state-of-the-art techniques and can learn to imitate behaviours from a single video demonstration.
Glen Berseth, Florian Golemo, Christopher Joseph Pal
J. Mach. Learn. Res.2
2023 Visual question answering from another perspective: CLEVR mental rotation tests
Christopher Beckham, Martin Weiss, Florian Golemo, Sina Honari, Derek Nowrouzezahrai, Christopher Joseph Pal
Pattern Recognit.3
2022 Kubric: A scalable dataset generator
abstract
Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and training details. But collecting, processing and annotating real data at scale is difficult, expensive, and frequently raises additional privacy, fairness and legal concerns. Synthetic data is a powerful tool with the potential to address these shortcomings: 1) it is cheap 2) supports rich ground-truth annotations 3) offers full control over data and 4) can circumvent or mitigate problems regarding bias, privacy and licensing. Unfortunately, software tools for effective data generation are less mature than those for architecture design and training, which leads to fragmented generation efforts. To address these problems we introduce Kubric, an open-source Python framework that interfaces with PyBullet and Blender to generate photo-realistic scenes, with rich annotations, and seamlessly scales to large jobs distributed over thousands of machines, and generating TBs of data. We demonstrate the effectiveness of Kubric by presenting a series of 13 different generated datasets for tasks ranging from studying 3D NeRF models to optical flow estimation. We release Kubric, the used assets, all of the generation code, as well as the rendered datasets for reuse and modification.
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J. Fleet, Dan Gnanapragasam, Florian Golemo, Charles Herrmann, Thomas Kipf, Abhijit Kundu, Dmitry Lagun, Issam H. Laradji, Hsueh-Ti Derek Liu, Henning Meyer, Yishu Miao, Derek Nowrouzezahrai, A. Cengiz Öztireli, Etienne Pot, Noha Radwan, Daniel Rebain, Sara Sabour, Mehdi S. M. Sajjadi, Matan Sela, Vincent Sitzmann, Austin Stone, Deqing Sun, Suhani Vora, Tianhao Wu 0003, Kwang Moo Yi, Fangcheng Zhong, Andrea Tagliasacchi
CVPR9
2022 Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction
Roger Girgis, Florian Golemo, Felipe Codevilla, Martin Weiss, Jim Aldon D'Souza, Samira Ebrahimi Kahou, Felix Heide, Christopher Joseph Pal
ICLR2
2021 gradSim: Differentiable simulation for system identification and visuomotor control
Krishna Murthy Jatavallabhula, Miles Macklin, Florian Golemo, Vikram Voleti, Linda Petrini, Martin Weiss, Breandan Considine, Jérôme Parent-Lévesque, Kevin Xie, Kenny Erleben, Liam Paull, Florian Shkurti, Derek Nowrouzezahrai, Sanja Fidler
ICLR3
2021 Sim2Real in Robotics and Automation: Applications and Challenges
abstract
To Perform reliably and consistently over sustained periods of time, large-scale automation critically relies on computer simulation. Simulation allows us and supervisory AI to effectively design, validate, and continuously improve complex processes, and helps practitioners to gain insight into the operation and justify future investments. While numerous successful applications of simulation in industry exist, such as circuit simulation, finite element methods, and computeraided design (CAD), state-of-the-art simulators fall short of accurately modeling physical phenomena, such as friction, impact, and deformation.
Sebastian Höfer, Kostas E. Bekris, Ankur Handa, Juan Camilo Gamboa, Melissa Mozifian, Florian Golemo, Christopher G. Atkeson, Dieter Fox, Kenneth Y. Goldberg, John J. Leonard, C. Karen Liu, Jan Peters 0001, Shuran Song, Peter Welinder, Martha White
IEEE Trans Autom. Sci. Eng.6
2020 Unsupervised Learning of Dense Visual Representations
abstract
Contrastive self-supervised learning has emerged as a promising approach to unsupervised visual representation learning. In general, these methods learn global (image-level) representations that are invariant to different views (i.e., compositions of data augmentation) of the same image. However, many visual understanding tasks require dense (pixel-level) representations. In this paper, we propose View-Agnostic Dense Representation (VADeR) for unsupervised learning of dense representations. VADeR learns pixelwise representations by forcing local features to remain constant over different viewing conditions. Specifically, this is achieved through pixel-level contrastive learning: matching features (that is, features that describes the same location of the scene on different views) should be close in an embedding space, while non-matching features should be apart. VADeR provides a natural representation for dense prediction tasks and transfers well to downstream tasks. Our method outperforms ImageNet supervised pretraining (and strong unsupervised baselines) in multiple dense prediction tasks.
Pedro O. Pinheiro, Amjad Almahairi, Ryan Y. Benmalek, Florian Golemo, Aaron C. Courville
NeurIPS4
2020 Pix2Shape: Towards Unsupervised Learning of 3D Scenes from Images Using a View-Based Representation
Sai Rajeswar, Fahim Mannan, Florian Golemo, Jérôme Parent-Lévesque, David Vázquez 0001, Derek Nowrouzezahrai, Aaron C. Courville
Int. J. Comput. Vis.3
2017 A multimodal dataset for object model learning from natural human-robot interaction
abstract
Learning object models in the wild from natural human interactions is an essential ability for robots to perform general tasks. In this paper we present a robocentric multimodal dataset addressing this key challenge. Our dataset focuses on interactions where the user teaches new objects to the robot in various ways. It contains synchronized recordings of visual (3 cameras) and audio data which provide a challenging evaluation framework for different tasks. Additionally, we present an end-to-end system that learns object models using object patches extracted from the recorded natural interactions. Our proposed pipeline follows these steps: (a) recognizing the interaction type, (b) detecting the object that the interaction is focusing on, and (c) learning the models from the extracted data. Our main contribution lies in the steps towards identifying the target object patches of the images. We demonstrate the advantages of combining language and visual features for the interaction recognition and use multiple views to improve the object modelling. Our experimental results show that our dataset is challenging due to occlusions and domain change with respect to typical object learning frameworks. The performance of common out-of-the-box classifiers trained on our data is low. We demonstrate that our algorithm outperforms such baselines.
Pablo Azagra, Florian Golemo, Yoan Mollard, Manuel Lopes 0001, Javier Civera 0001, Ana Cristina Murillo
IROS2