Oleh Kolner

dblp:388/2516 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 61% Robot navigation and mapping · 30% Efficient and distributed learning · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
active perception
0.912025
Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning · ICLR 2025
Computer vision › Vision and language
visual reasoning
0.912025
Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning · ICLR 2025
Computer vision › Vision and language › visual reasoning
visual relationship reasoning
0.912025
Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning · ICLR 2025
Machine learning › Efficient and distributed learning
data-efficient learning
0.312025
Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning · ICLR 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.9active vision · 0.9
YearPublicationVenuePosition
2025 Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning
abstract
Human capabilities in understanding visual relations are far superior to those of AI systems, especially for previously unseen objects. For example, while AI systems struggle to determine whether two such objects are visually the same or different, humans can do so with ease. Active vision theories postulate that the learning of visual relations is grounded in actions that we take to fixate objects and their parts by moving our eyes. In particular, the low-dimensional spatial information about the corresponding eye movements is hypothesized to facilitate the representation of relations between different image parts. Inspired by these theories, we develop a system equipped with a novel Glimpse-based Active Perception (GAP) that sequentially glimpses at the most salient regions of the input image and processes them at high resolution. Importantly, our system leverages the locations stemming from the glimpsing actions, along with the visual content around them, to represent relations between different parts of the image. The results suggest that the GAP is essential for extracting visual relations that go beyond the immediate visual content. Our approach reaches state-of-the-art performance on several visual reasoning tasks being more sample-efficient, and generalizing better to out-of-distribution visual inputs than prior models.
Oleh Kolner, Thomas Ortner, Stanislaw Wozniak, Angeliki Pantazi
ICLR1