EDBT 2026 Demo / reviewers in the wild / expert
Matjaz Jogan
dblp:44/1083
· DBLP profile ↗
12ranked-venue papers
5as first author
4since 2021 · last 2026
0000-0003-3771-3146ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 2 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Representation and self-supervised learning · 42% Efficient and distributed learning · 42% Transfer learning and domain adaptation · 6% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
0.9 | 1 | 2025 | FORLA: Federated Object-Centric Representation Learning with Slot Attention · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › federated learning
federated representation learning |
0.9 | 1 | 2025 | FORLA: Federated Object-Centric Representation Learning with Slot Attention · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning |
0.9 | 1 | 2025 | FORLA: Federated Object-Centric Representation Learning with Slot Attention · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning › object-centric representation learning
slot attention |
0.9 | 1 | 2025 | FORLA: Federated Object-Centric Representation Learning with Slot Attention · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
feature adaptation |
0.3 | 1 | 2025 | FORLA: Federated Object-Centric Representation Learning with Slot Attention · NeurIPS 2025 |
Computer vision › Video understanding and tracking
cue integration |
0.2 | 1 | 2013 | Optimal integration of visual speed across different spatiotemporal frequency channels · NIPS 2013 |
Image and video processing
image representation |
0.0 | 1 | 2003 | Karhunen-Loéve expansion of a set of rotated templates · IEEE Trans. Image Process. 2003 |
Image and video processing › subspace analysis › principal component analysis
karhunen-loeve transform |
0.0 | 1 | 2003 | Karhunen-Loéve expansion of a set of rotated templates · IEEE Trans. Image Process. 2003 |
Robotics › Robot navigation and mapping › localization
appearance-based localization |
0.0 | 1 | 2002 | Mobile Robot Localization using an Incremental Eigenspace Model · ICRA 2002 |
Robotics › Robot navigation and mapping
localization |
0.0 | 1 | 2002 | Mobile Robot Localization using an Incremental Eigenspace Model · ICRA 2002 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
online subspace learning |
0.0 | 1 | 2002 | Mobile Robot Localization using an Incremental Eigenspace Model · ICRA 2002 |
Methods — techniques the papers use, named apart from their topics
student-teacher architecture · 0.9slot attention · 0.9foundation model adaptation · 0.9psychophysics · 0.2divisive normalization · 0.2bayesian inference · 0.2karhunen-loéve expansion · 0.0covariance matrix decomposition · 0.0incremental PCA · 0.0eigenspace model · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Slot-BERT: Self-supervised object discovery in surgical videoabstract• Introduce a novel object-centric self-supervised representation learning model based on bidirectional temporal reasoning across video frames. • Introduce slot-contrastive loss, specifically designed for slot attention, to improve orthogonality between slots. • Superior temporal coherence and zero-shot generalization across four surgical video datasets from three different domains: abdominal, cholecystectomy, and thoracic surgery. Object-centric slot attention is a powerful framework for unsupervised learning of structured and explainable representations that can support reasoning about objects and actions, including in surgical video. However, current object-centric models either fail to reliably capture object dependencies in seconds-long video episodes that encompass surgical actions and tasks or are computationally too expensive for practical implementation. We introduce Slot-BERT, a slot attention model with a temporal slot transformer module to overcome these limitations. Our core innovations are: 1) A bidirectional transformer module that processes object-centric slot representations, enabling longer-range temporal coherence; 2) A slot-contrastive loss that further improves the representation by enforcing slot dissimilarity; 3) We evaluate Slot-BERT on real-world surgical video datasets from abdominal, cholecystectomy, and thoracic procedures, and on real and synthetic videos with everyday objects. Our method surpasses state-of-the-art object-centric approaches under unsupervised training achieving superior performance across these domains. We also demonstrate efficient zero-shot domain adaptation to data from diverse surgical specialties and databases. Guiqiu Liao, Matjaz Jogan, Marcel Hussing, Kenta Nakahashi, Kazuhiro Yasufuku, Amin Madani, Eric Eaton, Daniel A. Hashimoto |
Medical Image Anal. | 2 |
| 2025 | Future Slot Prediction for Unsupervised Object Discovery in Surgical Video
Guiqiu Liao, Matjaz Jogan, Marcel Hussing, Eric Eaton, Daniel A. Hashimoto |
MICCAI (11) | 2 |
| 2025 | FORLA: Federated Object-Centric Representation Learning with Slot AttentionabstractLearning efficient visual representations across heterogeneous unlabeled datasets remains a central challenge in federated learning. Effective federated representations require features that are jointly informative across clients while disentangling domain-specific factors without supervision. We introduce FORLA, a novel framework for federated object-centric representation learning and feature adaptation across clients using unsupervised slot attention. At the core of our method is a shared feature adapter, trained collaboratively across clients to adapt features from foundation models, and a shared slot attention module that learns to reconstruct the adapted features. To optimize this adapter, we design a two-branch student–teacher architecture. In each client, a student decoder learns to reconstruct full features from foundation models, while a teacher decoder reconstructs their adapted, low-dimensional counterpart. The shared slot attention module bridges cross-domain learning by aligning object-level representations across clients. Experiments in multiple real-world datasets show that our framework not only outperforms centralized baselines on object discovery but also learns a compact, universal representation that generalizes well across domains. This work highlights federated slot attention as an effective tool for scalable, unsupervised visual representation learning from cross-domain data with distributed concepts. Guiqiu Liao, Matjaz Jogan, Eric Eaton, Daniel A. Hashimoto |
NeurIPS | 2 |
| 2025 | Disentangling Spatio-Temporal Knowledge for Weakly Supervised Object Detection and Segmentation in Surgical VideoabstractWeakly supervised video object segmentation (WSVOS) enables the identification of segmentation maps without requiring extensive annotations of object masks, relying instead on coarse video labels indicating object presence. WSVOS in surgical videos is, however, more challenging due to the complex interaction of multiple transient objects, such as surgical tools moving in and out of the surgical field. In this scenario, state-of-the-art WSVOS methods struggle to learn accurate segmentation maps. We address this problem by introducing ViDeo Spatio-Temporal disentanglement Networks (VDST-Net), a framework to disentangle complex spatio-temporal object interactions using semi-decoupled knowledge distillation to predict high-quality class activation maps (CAMs). A teacher network is designed to help a temporal-reasoning student network resolve activation conflicts, as the student leverages temporal dependencies when specifics about object location and timing in the video are not provided. We demonstrate the efficacy of our framework on a challenging surgical video dataset where objects are, on average, present in less than 60% of annotated frames, and compare our method to state-of-the-art methods on surgical data and on a public dataset commonly used to benchmark WSVOS. Our method outperforms state-of-the-art techniques and generates accurate segmentation masks under video-level weak supervision. Our code is available at: https://github.com/PCASOlab/VDST-net. Guiqiu Liao, Matjaz Jogan, Sai Koushik, Eric Eaton, Daniel A. Hashimoto |
WACV | 2 |
| 2013 | Optimal integration of visual speed across different spatiotemporal frequency channelsabstractHow does the human visual system compute the speed of a coherent motion stimulus that contains motion energy in different spatiotemporal frequency bands? Here we propose that perceived speed is the result of optimal integration of speed information from independent spatiotemporal frequency tuned channels. We formalize this hypothesis with a Bayesian observer model that treats the channel activity as independent cues, which are optimally combined with a prior expectation for slow speeds. We test the model against behavioral data from a 2AFC speed discrimination task with which we measured subjects' perceived speed of drifting sinusoidal gratings with different contrasts and spatial frequencies, and of various combinations of these single gratings. We find that perceived speed of the combined stimuli is independent of the relative phase of the underlying grating components, and that the perceptual biases and discrimination thresholds are always smaller for the combined stimuli, supporting the cue combination hypothesis. The proposed Bayesian model fits the data well, accounting for perceptual biases and thresholds of both simple and combined stimuli. Fits are improved if we assume that the channel responses are subject to divisive normalization, which is in line with physiological evidence. Our results provide an important step toward a more complete model of visual motion perception that can predict perceived speeds for stimuli of arbitrary spatial structure. Matjaz Jogan, Alan A. Stocker |
NIPS | 1 |
| 2008 | Unsupervised Learning of a Hierarchy of Topological Maps Using Omnidirectional ImagesabstractThis paper presents a novel appearance-based method for path-based map learning by a mobile robot equipped with an omnidirectional camera. In particular, we focus on an unsupervised construction of topological maps, which provide an abstraction of the environment in terms of visual aspects. An unsupervised clustering algorithm is used to represent the images in multiple subspaces, forming thus a sensory grounded representation of the environment's appearance. By introducing transitional fields between clusters we are able to obtain a partitioning of the image set into distinctive visual aspects. By abstracting the low-level sensory data we are able to efficiently reconstruct the overall topological layout of the covered path. After the high level topology is estimated, we repeat the procedure on the level of visual aspects to obtain local topological maps. We demonstrate how the resulting representation can be used for modeling indoor and outdoor environments, how it successfully detects previously visited locations and how it can be used for the estimation of the current visual aspect and the retrieval of the relative position within the current visual aspect. Ales Stimec, Matjaz Jogan, Ales Leonardis |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2005 | Panoramic volumes for robot localizationabstractWe propose a method for visual robot localization using a panoramic image volume as the representation from which we can generate views from virtual viewpoints and match them to the current view. We use a geometric image-based rendering formalism in combination with a subspace representation of images, which allows us to synthesize views at arbitrary virtual viewpoints from a compact low-dimensional representation. Matej Artac, Matjaz Jogan, Ales Leonardis, Hynek Bakstein |
IROS | 2 |
| 2003 | A Framework for Robust and Incremental Self-Localization of a Mobile Robot
Matjaz Jogan, Matej Artac, Danijel Skocaj, Ales Leonardis |
ICVS | 1 |
| 2003 | Karhunen-Loéve expansion of a set of rotated templatesabstractIn this paper, we propose a novel method for efficiently calculating the eigenvectors of uniformly rotated images of a set of templates. As we show, the images can be optimally approximated by a linear series of eigenvectors which can be calculated without actually decomposing the sample covariance matrix. Matjaz Jogan, Emil Zagar, Ales Leonardis |
IEEE Trans. Image Process. | 1 |
| 2002 | Mobile Robot Localization using an Incremental Eigenspace ModelabstractWhen using appearance-based recognition for self-localization of mobile robots, the images obtained during the exploration of the environment need to be efficiently stored in the memory. PCA offers means for representing the images in a low-dimensional subspace, which allows for efficient matching and recognition. For active exploration it is necessary to use an incremental method for the computation of the subspace. We propose to use an incremental PCA algorithm with the updating of partial image representations in a way that allows the robot to discard the acquired images immediately after the update. Such a model is open-ended, meaning that we can easily update it with new images. We show that the performance of the proposed method is comparable to the performance of the batch method in terms of compression, computational cost and the precision of localization. We also show that by applying the repetitive learning, the subspace converges to that constructed with the batch method. Matej Artac, Matjaz Jogan, Ales Leonardis |
ICRA | 2 |
| 2000 | Robust Localization Using Panoramic View-Based RecognitionabstractThe results of earlier studies on the possibility of spatial localization from panoramic images have shown good prospects for view-based methods. The major advantages of these methods are a wide field-of-view, capability of modeling cluttered environments, and flexibility in the learning phase. The redundant information captured in similar views is efficiently handled by the eigenspace approach. However, the standard approaches are sensitive to noise and occlusion. We present a method of view-based localization in a robust framework that solves these problems to a large degree. Experimental results on a large set of real panoramic images demonstrate the effectiveness of the approach and the level of achieved robustness. Matjaz Jogan, Ales Leonardis |
ICPR | 1 |
| 1999 | Panoramic Eigenimages for Spatial Localisation
Matjaz Jogan, Ales Leonardis |
CAIP | 1 |