Aleksis Pirinen

dblp:191/4639 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 26% Robot navigation and mapping · 16% Face, body and person analysis · 11%
Theoretical computer science
1 paper
Mathematical optimization · 87% Graph algorithms and graph theory · 13%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 25 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
deep reinforcement learning
1.332021
Embodied Visual Active Learning for Semantic Segmentation · AAAI 2021
Deep Reinforcement Learning for Active Human Pose Estimation · AAAI 2020
Deep Reinforcement Learning of Region Proposal Networks for Object Detection · CVPR 2018
Computer vision › Face, body and person analysis
human pose estimation
0.822020
Deep Reinforcement Learning for Active Human Pose Estimation · AAAI 2020
Domes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction · NeurIPS 2019
Machine learning › Reinforcement learning › bandit › pure-exploration bandit
active search
0.812024
GOMAA-Geo: GOal Modality Agnostic Active Geo-localization · NeurIPS 2024
Robotics › Robot navigation and mapping › mobile robot navigation › 3d navigation
aerial robot navigation
0.812024
GOMAA-Geo: GOal Modality Agnostic Active Geo-localization · NeurIPS 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
GOMAA-Geo: GOal Modality Agnostic Active Geo-localization · NeurIPS 2024
Computer vision › Vision and language
cross-modal alignment
0.812024
GOMAA-Geo: GOal Modality Agnostic Active Geo-localization · NeurIPS 2024
Robotics › Legged, aerial and field robots › field robotics
search and rescue
0.812024
GOMAA-Geo: GOal Modality Agnostic Active Geo-localization · NeurIPS 2024
Machine learning › Reinforcement learning
policy learning
0.622019
Domes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction · NeurIPS 2019
Reinforcement Learning for Visual Object Detection · CVPR 2016
Computer vision › Image recognition and object detection
object detection
0.622018
Deep Reinforcement Learning of Region Proposal Networks for Object Detection · CVPR 2018
Reinforcement Learning for Visual Object Detection · CVPR 2016
Robotics › Robot navigation and mapping › view planning
next-best-view planning
0.522020
Domes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction · NeurIPS 2019
Deep Reinforcement Learning for Active Human Pose Estimation · AAAI 2020
Robotics › Robot navigation and mapping › active perception
active exploration
0.512021
Embodied Visual Active Learning for Semantic Segmentation · AAAI 2021
Computer vision › Segmentation and scene understanding › annotation-efficient segmentation
active learning for segmentation
0.512021
Embodied Visual Active Learning for Semantic Segmentation · AAAI 2021
Knowledge, reasoning and agents › Multi-agent systems › autonomous agents
embodied agent
0.512021
Embodied Visual Active Learning for Semantic Segmentation · AAAI 2021
Computer vision › Segmentation and scene understanding
semantic segmentation
0.512021
Embodied Visual Active Learning for Semantic Segmentation · AAAI 2021
Computer vision › 3D vision
3d human pose estimation
0.412019
Domes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction · NeurIPS 2019
Computer vision › Face, body and person analysis › human pose estimation
3d pose estimation
0.412019
Domes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction · NeurIPS 2019
Data mining
clustering
0.412019
Exact Clustering of Weighted Graphs via Semidefinite Programming · J. Mach. Learn. Res. 2019
Data mining › clustering
graph clustering
0.412019
Exact Clustering of Weighted Graphs via Semidefinite Programming · J. Mach. Learn. Res. 2019
Mathematical optimization
semidefinite programming
0.412019
Exact Clustering of Weighted Graphs via Semidefinite Programming · J. Mach. Learn. Res. 2019
Mathematical optimization › convex relaxation
semidefinite relaxation
0.412019
Exact Clustering of Weighted Graphs via Semidefinite Programming · J. Mach. Learn. Res. 2019
Computer vision › Image recognition and object detection › object detection › object proposal generation
region proposal network
0.312018
Deep Reinforcement Learning of Region Proposal Networks for Object Detection · CVPR 2018
Machine learning › Reinforcement learning › policy learning
search policy learning
0.212016
Reinforcement Learning for Visual Object Detection · CVPR 2016
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
sequential search
0.212016
Reinforcement Learning for Visual Object Detection · CVPR 2016
Graph algorithms and graph theory
graph clustering
0.112019
Exact Clustering of Weighted Graphs via Semidefinite Programming · J. Mach. Learn. Res. 2019
Computer vision › Image recognition and object detection › object detection
sliding window detection
0.112016
Reinforcement Learning for Visual Object Detection · CVPR 2016

Methods — techniques the papers use, named apart from their topics

deep reinforcement learning · 1.2reinforcement learning · 1.0semidefinite programming · 0.8foundation model pretraining · 0.8contrastive learning · 0.8online retraining · 0.5annotation propagation · 0.5multi-view fusion · 0.4monocular pose estimation · 0.4self-supervised learning · 0.4active triangulation · 0.4
YearPublicationVenuePosition
2024 GOMAA-Geo: GOal Modality Agnostic Active Geo-localization
abstract
We consider the task of active geo-localization (AGL) in which an agent uses a sequence of visual cues observed during aerial navigation to find a target specified through multiple possible modalities. This could emulate a UAV involved in a search-and-rescue operation navigating through an area, observing a stream of aerial images as it goes. The AGL task is associated with two important challenges. Firstly, an agent must deal with a goal specification in one of multiple modalities (e.g., through a natural language description) while the search cues are provided in other modalities (aerial imagery). The second challenge is limited localization time (e.g., limited battery life, urgency) so that the goal must be localized as efficiently as possible, i.e. the agent must effectively leverage its sequentially observed aerial views when searching for the goal. To address these challenges, we propose GOMAA-Geo -- a goal modality agnostic active geo-localization agent -- for zero-shot generalization between different goal modalities. Our approach combines cross-modality contrastive learning to align representations across modalities with supervised foundation model pretraining and reinforcement learning to obtain highly effective navigation and localization policies. Through extensive evaluations, we show that GOMAA-Geo outperforms alternative learnable approaches and that it generalizes across datasets -- e.g., to disaster-hit areas without seeing a single disaster scenario during training -- and goal modalities -- e.g., to ground-level imagery or textual descriptions, despite only being trained with goals specified as aerial views. Our code is available at: https://github.com/mvrl/GOMAA-Geo.
Anindya Sarkar, Srikumar Sastry, Aleksis Pirinen, Chongjie Zhang, Nathan Jacobs, Yevgeniy Vorobeychik
NeurIPS3
2021 Embodied Visual Active Learning for Semantic Segmentation
abstract
We study the task of embodied visual active learning, where an agent is set to explore a 3d environment with the goal to acquire visual scene understanding by actively selecting views for which to request annotation. While accurate on some benchmarks, today's deep visual recognition pipelines tend to not generalize well in certain real-world scenarios, or for unusual viewpoints. Robotic perception, in turn, requires the capability to refine the recognition capabilities for the conditions where the mobile system operates, including cluttered indoor environments or poor illumination. This motivates the proposed task, where an agent is placed in a novel environment with the objective of improving its visual recognition capability. To study embodied visual active learning, we develop a battery of agents - both learnt and pre-specified - and with different levels of knowledge of the environment. The agents are equipped with a semantic segmentation network and seek to acquire informative views, move and explore in order to propagate annotations in the neighbourhood of those views, then refine the underlying segmentation network by online retraining. The trainable method uses deep reinforcement learning with a reward function that balances two competing objectives: task performance, represented as visual recognition accuracy, which requires exploring the environment, and the necessary amount of annotated data requested during active exploration. We extensively evaluate the proposed models using the photorealistic Matterport3D simulator and show that a fully learnt method outperforms comparable pre-specified counterparts, even when requesting fewer annotations.
David Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian Sminchisescu
AAAI2
2020 Deep Reinforcement Learning for Active Human Pose Estimation
abstract
Most 3d human pose estimation methods assume that input – be it images of a scene collected from one or several viewpoints, or from a video – is given. Consequently, they focus on estimates leveraging prior knowledge and measurement by fusing information spatially and/or temporally, whenever available. In this paper we address the problem of an active observer with freedom to move and explore the scene spatially – in ‘time-freeze’ mode – and/or temporally, by selecting informative viewpoints that improve its estimation accuracy. Towards this end, we introduce Pose-DRL, a fully trainable deep reinforcement learning-based active pose estimation architecture which learns to select appropriate views, in space and time, to feed an underlying monocular pose estimator. We evaluate our model using single- and multi-target estimators with strong result in both settings. Our system further learns automatic stopping conditions in time and transition functions to the next temporal processing step in videos. In extensive experiments with the Panoptic multi-view setup, and for complex scenes containing multiple people, we show that our model learns to select viewpoints that yield significantly more accurate pose estimates compared to strong multi-view baselines.
Erik Gärtner, Aleksis Pirinen, Cristian Sminchisescu
AAAI2
2020 Semantic Synthesis of Pedestrian Locomotion
Maria Priisalu, Ciprian Paduraru, Aleksis Pirinen, Cristian Sminchisescu
ACCV (2)3
2019 Domes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction
abstract
Existing state-of-the-art estimation systems can detect 2d poses of multiple people in images quite reliably. In contrast, 3d pose estimation from a single image is ill-posed due to occlusion and depth ambiguities. Assuming access to multiple cameras, or given an active system able to position itself to observe the scene from multiple viewpoints, reconstructing 3d pose from 2d measurements becomes well-posed within the framework of standard multi-view geometry. Less clear is what is an informative set of viewpoints for accurate 3d reconstruction, particularly in complex scenes, where people are occluded by others or by scene objects. In order to address the view selection problem in a principled way, we here introduce ACTOR, an active triangulation agent for 3d human pose reconstruction. Our fully trainable agent consists of a 2d pose estimation network (any of which would work) and a deep reinforcement learning-based policy for camera viewpoint selection. The policy predicts observation viewpoints, the number of which varies adaptively depending on scene content, and the associated images are fed to an underlying pose estimator. Importantly, training the policy requires no annotations - given a 2d pose estimator, ACTOR is trained in a self-supervised manner. In extensive evaluations on complex multi-people scenes filmed in a Panoptic dome, under multiple viewpoints, we compare our active triangulation agent to strong multi-view baselines, and show that ACTOR produces significantly more accurate 3d pose reconstructions. We also provide a proof-of-concept experiment indicating the potential of connecting our view selection policy to a physical drone observer.
Aleksis Pirinen, Erik Gärtner, Cristian Sminchisescu
NeurIPS1
2019 Exact Clustering of Weighted Graphs via Semidefinite Programming
abstract
As a model problem for clustering, we consider the densest $k$-disjoint-clique problem of partitioning a weighted complete graph into $k$ disjoint subgraphs such that the sum of the densities of these subgraphs is maximized. We establish that such subgraphs can be recovered from the solution of a particular semidefinite relaxation with high probability if the input graph is sampled from a distribution of clusterable graphs. Specifically, the semidefinite relaxation is exact if the graph consists of \(k\) large disjoint subgraphs, corresponding to clusters, with weight concentrated within these subgraphs, plus a moderate number of nodes not belonging to any cluster. Further, we establish that if noise is weakly obscuring these clusters, i.e, the between-cluster edges are assigned very small weights, then we can recover significantly smaller clusters. For example, we show that in approximately sparse graphs, where the between-cluster weights tend to zero as the size $n$ of the graph tends to infinity, we can recover clusters of size polylogarithmic in $n$ under certain conditions on the distribution of edge weights. Empirical evidence from numerical simulations is also provided to support these theoretical phase transitions to perfect recovery of the cluster structure.
Aleksis Pirinen, Brendan P. W. Ames
J. Mach. Learn. Res.1
2018 Deep Reinforcement Learning of Region Proposal Networks for Object Detection
abstract
We propose drl-RPN, a deep reinforcement learning-based visual recognition model consisting of a sequential region proposal network (RPN) and an object detector. In contrast to typical RPNs, where candidate object regions (RoIs) are selected greedily via class-agnostic NMS, drl-RPN optimizes an objective closer to the final detection task. This is achieved by replacing the greedy RoI selection process with a sequential attention mechanism which is trained via deep reinforcement learning (RL). Our model is capable of accumulating class-specific evidence over time, potentially affecting subsequent proposals and classification scores, and we show that such context integration significantly boosts detection accuracy. Moreover, drl-RPN automatically decides when to stop the search process and has the benefit of being able to jointly learn the parameters of the policy and the detector, both represented as deep networks. Our model can further learn to search over a wide range of exploration-accuracy trade-offs making it possible to specify or adapt the exploration extent at test time. The resulting search trajectories are image- and category-dependent, yet rely only on a single policy over all object categories. Results on the MS COCO and PASCAL VOC challenges show that our approach outperforms established, typical state-of-the-art object detection pipelines.
Aleksis Pirinen, Cristian Sminchisescu
CVPR1
2016 Reinforcement Learning for Visual Object Detection
abstract
One of the most widely used strategies for visual object detection is based on exhaustive spatial hypothesis search. While methods like sliding windows have been successful and effective for many years, they are still brute-force, independent of the image content and the visual category being searched. In this paper we present principled sequential models that accumulate evidence collected at a small set of image locations in order to detect visual objects effectively. By formulating sequential search as reinforcement learning of the search policy (including the stopping condition), our fully trainable model can explicitly balance for each class, specifically, the conflicting goals of exploration - sampling more image regions for better accuracy -, and exploitation - stopping the search efficiently when sufficiently confident about the target's location. The methodology is general and applicable to any detector response function. We report encouraging results in the PASCAL VOC 2012 object detection test set showing that the proposed methodology achieves almost two orders of magnitude speed-up over sliding window methods.
Stefan Mathe, Aleksis Pirinen, Cristian Sminchisescu
CVPR2