Görkay Aydemir

dblp:323/7577 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0002-0315-3312ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Video understanding and tracking · 62% Autonomous driving · 19% Segmentation and scene understanding · 9%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
feature tracking
1.922026
Track-On2: Enhancing Online Point Tracking With Memory · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Track-On: Transformer-based Online Point Tracking with Memory · ICLR 2025
Computer vision › Video understanding and tracking › feature tracking
long-term point tracking
1.012026
Track-On2: Enhancing Online Point Tracking With Memory · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Video understanding and tracking › object tracking
online tracking
0.912025
Track-On: Transformer-based Online Point Tracking with Memory · ICLR 2025
Robotics › Autonomous driving › trajectory prediction
multi-agent trajectory prediction
0.712023
ADAPT: Efficient Multi-Agent Trajectory Prediction with Adaptation · ICCV 2023
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.712023
Self-supervised Object-Centric Learning for Videos · NeurIPS 2023
Robotics › Autonomous driving
trajectory prediction
0.712023
ADAPT: Efficient Multi-Agent Trajectory Prediction with Adaptation · ICCV 2023
Computer vision › Segmentation and scene understanding › object segmentation
unsupervised multi-object segmentation
0.712023
Self-supervised Object-Centric Learning for Videos · NeurIPS 2023
Computer vision › Video understanding and tracking
video object segmentation
0.712023
Self-supervised Object-Centric Learning for Videos · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

transformer · 1.9synthetic training · 1.0memory mechanism · 1.0spatial memory · 0.9context memory · 0.9slot attention · 0.7self-supervised learning · 0.7feature reconstruction · 0.7endpoint-conditioned prediction · 0.7dynamic weight learning · 0.7
YearPublicationVenuePosition
2026 Track-On2: Enhancing Online Point Tracking With Memory
abstract
In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across video frames under significant appearance changes, motion, and occlusion. We target the online setting, i.e., tracking points frame-by-frame, making it suitable for real-time and streaming applications. We extend our prior model Track-On into Track-On2, a simple and efficient transformer-based model for online long-term tracking. Track-On2 improves both performance and efficiency through architectural refinements, more effective use of memory, and improved synthetic training strategies. Unlike prior approaches that rely on full-sequence access or iterative updates, our model processes frames causally and maintains temporal coherence via a memory mechanism, which is key to handling drift and occlusions without requiring future frames. At inference, we perform coarse patch-level classification followed by refinement. Beyond architecture, we systematically study synthetic training setups and their impact on memory behavior, showing how they shape temporal robustness over long sequences. Through comprehensive experiments, Track-On2 achieves state-of-the-art results across five synthetic and real-world benchmarks, surpassing prior online trackers and even strong offline methods that exploit bidirectional context. These results highlight the effectiveness of causal, memory-based architectures trained purely on synthetic data as scalable solutions for real-world point tracking.
Görkay Aydemir, Weidi Xie, Fatma Güney
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Track-On: Transformer-based Online Point Tracking with Memory
abstract
In this paper, we consider the problem of long-term point tracking, which requires consistent identification of points across multiple frames in a video, despite changes in appearance, lighting, perspective, and occlusions. We target online tracking on a frame-by-frame basis, making it suitable for real-world, streaming scenarios. Specifically, we introduce Track-On, a simple transformer-based model designed for online long-term point tracking. Unlike prior methods that depend on full temporal modeling, our model processes video frames causally without access to future frames, leveraging two memory modules —spatial memory and context memory— to capture temporal information and maintain reliable point tracking over long time horizons. At inference time, it employs patch classification and refinement to identify correspondences and track points with high accuracy. Through extensive experiments, we demonstrate that Track-On sets a new state-of-the-art for online models and delivers superior or competitive results compared to offline approaches on seven datasets, including the TAP-Vid benchmark. Our method offers a robust and scalable solution for real-time tracking in diverse applications. Project page: https://kuis-ai.github.io/track_on
Görkay Aydemir, Xiongyi Cai, Weidi Xie, Fatma Güney
ICLR1
2023 ADAPT: Efficient Multi-Agent Trajectory Prediction with Adaptation
abstract
Forecasting future trajectories of agents in complex traffic scenes requires reliable and efficient predictions for all agents in the scene. However, existing methods for trajectory prediction are either inefficient or sacrifice accuracy. To address this challenge, we propose ADAPT, a novel approach for jointly predicting the trajectories of all agents in the scene with dynamic weight learning. Our approach outperforms state-of-the-art methods in both single-agent and multi-agent settings on the Argoverse and Interaction datasets, with a fraction of their computational overhead. We attribute the improvement in our performance: first, to the adaptive head augmenting the model capacity without increasing the model size; second, to our design choices in the endpoint-conditioned prediction, reinforced by gradient stopping. Our analyses show that ADAPT can focus on each agent with adaptive prediction, allowing for accurate predictions efficiently. https://KUIS-AI.github.io/adapt
Görkay Aydemir, Adil Kaan Akan, Fatma Güney
ICCV1
2023 Self-supervised Object-Centric Learning for Videos
abstract
Unsupervised multi-object segmentation has shown impressive results on images by utilizing powerful semantics learned from self-supervised pretraining. An additional modality such as depth or motion is often used to facilitate the segmentation in video sequences. However, the performance improvements observed in synthetic sequences, which rely on the robustness of an additional cue, do not translate to more challenging real-world scenarios. In this paper, we propose the first fully unsupervised method for segmenting multiple objects in real-world sequences. Our object-centric learning framework spatially binds objects to slots on each frame and then relates these slots across frames. From these temporally-aware slots, the training objective is to reconstruct the middle frame in a high-level semantic feature space. We propose a masking strategy by dropping a significant portion of tokens in the feature space for efficiency and regularization. Additionally, we address over-clustering by merging slots based on similarity. Our method can successfully segment multiple instances of complex and high-variety classes in YouTube videos.
Görkay Aydemir, Weidi Xie, Fatma Güney
NeurIPS1