Stephan Liwicki

dblp:60/9606 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
10since 2021 · last 2024
0000-0002-2002-3627ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 ReCoRe: Regularized Contrastive Representation Learning of World Model
abstract
While recent model-free Reinforcement Learning (RL) methods have demonstrated human-level effectiveness in gaming environments, their success in everyday tasks like visual navigation has been limited, particularly under significant appearance variations. This limitation arises from (i) poor sample efficiency and (ii) over-fitting to training scenarios. To address these challenges, we present a world model that learns invariant features using (i) contrastive unsupervised learning and (ii) an intervention-invariant regularizer. Learning an explicit representation of the world dynamics i.e. a world model, improves sample efficiency while contrastive learning implicitly enforces learning of invariant features, which improves generalization. However, the naïve integration of contrastive loss to world models is not good enough, as world-model-based RL methods independently optimize representation learning and agent policy. To overcome this issue, we propose an intervention-invariant regularizer in the form of an auxiliary task such as depth prediction, image denoising, image segmentation, etc., that explicitly enforces invariance to style interventions. Our method outperforms current state-of-the-art model-based and model-free RL methods and significantly improves on out-of-distribution point navigation tasks evaluated on the iGibson benchmark. With only visual observations, we further demonstrate that our approach outperforms recent language-guided foundation models for point navigation, which is essential for deployment on robots with limited computation capabilities. Finally, we demonstrate that our proposed model excels at the sim-to-real transfer of its perception module on the Gibson benchmark.
Rudra P. K. Poudel, Harit Pandya, Stephan Liwicki, Roberto Cipolla
CVPR3
2024 DiaLoc: An Iterative Approach to Embodied Dialog Localization
abstract
Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing dialogbased localization approaches assume the availability of entire dialog prior to Iocalizaiton, which is impractical for deployed dialog-based localization. In this paper, we propose DiaLoc, a new dialog-based localization framework which aligns with a real human operator behavior. Specifically, we produce an iterative refinement of location predictions which can visualize current pose believes after each dialog turn. DiaLoc effectively utilizes the multimodal data for multi-shot localization, where a fusion encoder fuses vision and dialog information iteratively. We achieve state-of-the-art results on embodied dialog-based localization task, in single-shot (+7.08% in Acc5@valUnseen) and multishot settings (+10.85% in Acc5@valUnseen). DiaLoc narrows the gap between simulation and real-world applications, opening doors for future research on collaborative localization and navigation.
Chao Zhang 0023, Mohan Li, Ignas Budvytis, Stephan Liwicki
CVPR4
2024 Recurrent Reinforcement Learning with Memoroids
abstract
Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially Observable Markov Decision Processes (POMDPs) by mapping trajectories to latent Markov states. Neither model scales particularly well to long sequences, especially compared to an emerging class of memory models called Linear Recurrent Models. We discover that the recurrent update of these models resembles a monoid, leading us to reformulate existing models using a novel monoid-based framework that we call memoroids. We revisit the traditional approach to batching in recurrent reinforcement learning, highlighting theoretical and empirical deficiencies. We leverage memoroids to propose a batching method that improves sample efficiency, increases the return, and simplifies the implementation of recurrent loss functions in reinforcement learning.
Steven D. Morad, Chris Lu 0001, Ryan Kortvelesy, Stephan Liwicki, Jakob N. Foerster, Amanda Prorok
NeurIPS4
2023 Hierarchical Quantization Consistency for Fully Unsupervised Image Retrieval
Guile Wu, Chao Zhang 0023, Stephan Liwicki
BMVC3
2023 POPGym: Benchmarking Partially Observable Reinforcement Learning
Steven D. Morad, Ryan Kortvelesy, Matteo Bettini, Stephan Liwicki, Amanda Prorok
ICLR4
2023 Reinforcement Learning with Fast and Forgetful Memory
abstract
Nearly all real world tasks are inherently partially observable, necessitating the use of memory in Reinforcement Learning (RL). Most model-free approaches summarize the trajectory into a latent Markov state using memory models borrowed from Supervised Learning (SL), even though RL tends to exhibit different training and efficiency characteristics. Addressing this discrepancy, we introduce Fast and Forgetful Memory, an algorithm-agnostic memory model designed specifically for RL. Our approach constrains the model search space via strong structural priors inspired by computational psychology. It is a drop-in replacement for recurrent neural networks (RNNs) in recurrent RL algorithms, achieving greater reward than RNNs across various recurrent benchmarks and algorithms _without changing any hyperparameters_. Moreover, Fast and Forgetful Memory exhibits training speeds two orders of magnitude faster than RNNs, attributed to its logarithmic time and linear space complexity. Our implementation is available at https://github.com/proroklab/ffm.
Steven D. Morad, Ryan Kortvelesy, Stephan Liwicki, Amanda Prorok
NeurIPS3
2023 HexNet: An Orientation-Aware Deep Learning Framework for Omni-Directional Input
abstract
While omni-directional sensors provide holistic representations typical deep learning frameworks reduce the benefits by introducing distortions and discontinuities as spherical data is supplied as planar input. On the other hand, recent spherical convolutional neural networks (CNNs) often require significant memory and parameters, thus enabling execution only at very low resolutions and shallow architectures. We propose HexNet, an orientation-aware deep learning framework for spherical signals, that allows for fast computation as we exploit standard planar network operations on an efficiently arranged projection of the sphere. Furthermore, we introduce a graph-based version for partial spheres, allowing us to compete at high-resolution with planar CNNs using residual network architectures. Our kernels operate on the tangent of the sphere and thus standard feature weights, pretrained on perspective data, can be transferred, enabling spherical pretraining on ImageNet. As our design is free of distortions and discontinuity, our orientation-aware CNN becomes a new state of the art for semantic segmentation on the recent 2D3DS dataset, and the omni-directional version of SYNTHIA introduced in this work. Moreover, we experimentally show the benefit of our spherical representation over standard images on the Cityscapes dataset by reducing distortion effects of planar CNNs. We implement object detection for the spherical domain. Rotation invariant classification and segmentation tasks are additionally presented for comparison to prior art.
Chao Zhang 0023, Stephan Liwicki, Sen He 0001, William A. P. Smith, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Beyond the CLS Token: Image Reranking using Pretrained Vision Transformers
Chao Zhang 0023, Stephan Liwicki, Roberto Cipolla
BMVC2
2022 CoMBiNED: Multi-Constrained Model Based Planning for Navigation in Dynamic Environments
abstract
Recent deep reinforcement learning (DRL) approaches have achieved high success rate in map-less dynamic obstacle avoidance tasks. However, navigation in unseen dynamic scenarios without a pre-built map in the presence of dynamic obstacles still remains an open challenge. Since, learning accurate models for complex robotic scenarios such as navigation directly from high dimensional sensory measurements requires a large amount of data and training. Furthermore, even a small change on robot configuration such as kino-dynamics or sensor in the inference time requires re-training of the policy. In this paper, we address these issues in a principled fashion through a multi-constraint model based online planning (CoMBiNED) framework that does not require any retraining or modifications on the existing policy. We disentangle the given task into sub-tasks and learn dynamical models for them. Treating these dynamical models as soft-constraints, we employ stochastic optimisation to jointly optimize these sub-tasks on-the-fly at the inference time. We consider navigation as central application in this work and evaluate our approach on publicly available benchmark with complex dynamic scenarios and achieved significant improvement over recent approaches both in the cases of with-and-without given map of the environment.
Harit Pandya, Rudra P. K. Poudel, Stephan Liwicki
IROS3
2021 Lifted Semantic Graph Embedding for Omnidirectional Place Recognition
abstract
Typical place recognition is dependent on the visual appearance and camera position of query images, without explicit use of domain knowledge and geometric relationships between key features in the scene. We exploit semantic grouping of pixels, and camera-pose robust scene graphs to perform structure-based visual localization for place recognition. In particular, we first formulate place recognition as an image retrieval task. Then, we lift the omnidirectional input images into 3D space, and compute a rotation and translation invariant semantic graph embedding to encode query and reference images. Finally, place information is obtained through graph similarity matching. Our graph representation is a simple addition to standard image embeddings with minimal overhead, but contains awareness of objects and their geometric relationships. In our experiments, we show improvement over typical place recognition, especially in environments with repetitions and dynamic appearance changes.
Chao Zhang 0023, Ignas Budvytis, Stephan Liwicki, Roberto Cipolla
3DV3
2020 Rotation Equivariant Orientation Estimation for Omnidirectional Localization
Chao Zhang 0023, Ignas Budvytis, Stephan Liwicki, Roberto Cipolla
ACCV (4)3
2020 A Spherical Approach to Planar Semantic Segmentation
Chao Zhang 0023, Sen He 0001, Stephan Liwicki
BMVC3
2019 Fast-SCNN: Fast Semantic Segmentation Network
Rudra P. K. Poudel, Stephan Liwicki, Roberto Cipolla
BMVC2
2019 Orientation-Aware Semantic Segmentation on Icosahedron Spheres
abstract
We address semantic segmentation on omnidirectional images, to leverage a holistic understanding of the surrounding scene for applications like autonomous driving systems. For the spherical domain, several methods recently adopt an icosahedron mesh, but systems are typically rotation invariant or require significant memory and parameters, thus enabling execution only at very low resolutions. In our work, we propose an orientation-aware CNN framework for the icosahedron mesh. Our representation allows for fast network operations, as our design simplifies to standard network operations of classical CNNs, but under consideration of north-aligned kernel convolutions for features on the sphere. We implement our representation and demonstrate its memory efficiency up-to a level-8 resolution mesh (equivalent to 640 x 1024 equirectangular images). Finally, since our kernels operate on the tangent of the sphere, standard feature weights, pretrained on perspective data, can be directly transferred with only small need for weight refinement. In our evaluation our orientation-aware CNN becomes a new state of the art for the recent 2D3DS dataset, and our Omni-SYNTHIA version of SYNTHIA. Rotation invariant classification and segmentation tasks are additionally presented for comparison to prior art.
Chao Zhang 0023, Stephan Liwicki, William A. P. Smith, Roberto Cipolla
ICCV2
2018 ContextNet: Exploring Context and Detail for Semantic Segmentation in Real-time
Rudra P. K. Poudel, Ujwal Bonde, Stephan Liwicki, Christopher Zach
BMVC3
2017 Scale Exploiting Minimal Solvers for Relative Pose with Calibrated Cameras
Stephan Liwicki, Christopher Zach
BMVC1
2016 Coarse-to-fine Planar Regularization for Dense Monocular Depth Estimation
Stephan Liwicki, Christopher Zach, Ondrej Miksik, Philip Torr 0001
ECCV (2)1
2015 Online Kernel Slow Feature Analysis for Temporal Video Segmentation and Tracking
abstract
Slow feature analysis (SFA) is a dimensionality reduction technique which has been linked to how visual brain cells work. In recent years, the SFA was adopted for computer vision tasks. In this paper, we propose an exact kernel SFA (KSFA) framework for positive definite and indefinite kernels in Krein space. We then formulate an online KSFA which employs a reduced set expansion. Finally, by utilizing a special kind of kernel family, we formulate exact online KSFA for which no reduced set is required. We apply the proposed system to develop a SFA-based change detection algorithm for stream data. This framework is employed for temporal video segmentation and tracking. We test our setup on synthetic and real data streams. When combined with an online learning tracking system, the proposed change detection approach improves upon tracking setups that do not utilize change detection.
Stephan Liwicki, Stefanos Zafeiriou, Maja Pantic
IEEE Trans. Image Process.1
2014 Full-Angle Quaternions for Robustly Matching Vectors of 3D Rotations
abstract
In this paper we introduce a new distance for robustly matching vectors of 3D rotations. A special representation of 3D rotations, which we coin full-angle quaternion (FAQ), allows us to express this distance as Euclidean. We apply the distance to the problems of 3D shape recognition from point clouds and 2D object tracking in color video. For the former, we introduce a hashing scheme for scale and translation which outperforms the previous state-of-the-art approach on a public dataset. For the latter, we incorporate online subspace learning with the proposed FAQ representation to highlight the benefits of the new representation.
Stephan Liwicki, Minh-Tri Pham, Stefanos Zafeiriou, Maja Pantic, Björn Stenger
CVPR1
2013 Euler Principal Component Analysis
abstract
Principal Component Analysis (PCA) is perhaps the most prominent learning tool for dimensionality reduction in pattern recognition and computer vision. However, the ℓ 2-norm employed by standard PCA is not robust to outliers. In this paper, we propose a kernel PCA method for fast and robust PCA, which we call Euler-PCA (e-PCA). In particular, our algorithm utilizes a robust dissimilarity measure based on the Euler representation of complex numbers. We show that Euler-PCA retains PCA’s desirable properties while suppressing outliers. Moreover, we formulate Euler-PCA in an incremental learning framework which allows for efficient computation. In our experiments we apply Euler-PCA to three different computer vision applications for which our method performs comparably with other state-of-the-art approaches.
Stephan Liwicki, Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic
Int. J. Comput. Vis.1
2012 Incremental Slow Feature Analysis with Indefinite Kernel for Online Temporal Video Segmentation
Stephan Liwicki, Stefanos Zafeiriou, Maja Pantic
ACCV (2)1
2012 Efficient Online Subspace Learning With an Indefinite Kernel for Visual Tracking and Recognition
abstract
We propose an exact framework for online learning with a family of indefinite (not positive) kernels. As we study the case of nonpositive kernels, we first show how to extend kernel principal component analysis (KPCA) from a reproducing kernel Hilbert space to Krein space. We then formulate an incremental KPCA in Krein space that does not require the calculation of preimages and therefore is both efficient and exact. Our approach has been motivated by the application of visual tracking for which we wish to employ a robust gradient-based kernel. We use the proposed nonlinear appearance model learned online via KPCA in Krein space for visual tracking in many popular and difficult tracking scenarios. We also show applications of our kernel framework for the problem of face recognition.
Stephan Liwicki, Stefanos Zafeiriou, Georgios Tzimiropoulos, Maja Pantic
IEEE Trans. Neural Networks Learn. Syst.1
2011 Fast and robust appearance-based tracking
abstract
We introduce a fast and robust subspace-based approach to appearance-based object tracking. The core of our approach is based on Fast Robust Correlation (FRC), a recently proposed technique for the robust estimation of large translational displacements. We show how the basic principles of FRC can be naturally extended to formulate a robust version of Principal Component Analysis (PCA) which can be efficiently implemented incrementally and therefore is particularly suitable for robust real-time appearance-based object tracking. Our experimental results demonstrate that the proposed approach outperforms other state-of-the-art holistic appearance-based trackers on several popular video sequences.
Stephan Liwicki, Stefanos Zafeiriou, Georgios Tzimiropoulos, Maja Pantic
FG1