Tim Verbelen

dblp:71/8853 · DBLP profile ↗
← Back
33ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0003-2731-7262ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 14 since 2021Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Representing Positional Information in Generative World Models for Object Manipulation
abstract
Object manipulation capabilities are essential skills that set apart embodied agents engaging with the world, especially in the realm of robotics. The ability to predict outcomes of interactions with objects is paramount in this setting. Although model-based control methods have started to be employed to tackle manipulation tasks, they have faced challenges in accurately manipulating objects. As we analyze the causes of this limitation, we identify the cause of underperformance in the way current world models represent crucial positional information, especially about the target’s goal specification for object positioning tasks. We introduce a general approach that empowers world model-based agents to effectively solve object positioning tasks. We propose two declinations of this approach for generative world models: position-conditioned (PCP) and latent-conditioned (LCP) policy learning. In particular, LCP employs object-centric latent representations that explicitly capture object positional information for goal specification. This naturally leads to the emergence of multimodal capabilities, enabling the specification of goals through spatial coordinates or a visual goal. Our methods are rigorously evaluated across several manipulation environments, showing favorable performance compared to current model-based control approaches.
Stefano Ferraro, Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Sai Rajeswar
ECAI3
2025 Active Inference and Intentional Behavior
abstract
Recent advances in theoretical biology suggest that key definitions of basal cognition and sentient behavior may arise as emergent properties of in vitro cell cultures and neuronal networks. Such neuronal networks reorganize activity to demonstrate structured behaviors when embodied in structured information landscapes. In this article, we characterize this kind of self-organization through the lens of the free energy principle, that is, as self-evidencing. We do this by first discussing the definitions of reactive and sentient behavior in the setting of active inference, which describes the behavior of agents that model the consequences of their actions. We then introduce a formal account of intentional behavior that describes agents as driven by a preferred end point or goal in latent state-spaces. We then investigate these forms of (reactive, sentient, and intentional) behavior using simulations. First, we simulate the in vitro experiments, in which neuronal cultures modulated activity to improve gameplay in a simplified version of Pong by implementing nested, free energy minimizing processes. The simulations are then used to deconstruct the ensuing predictive behavior, leading to the distinction between merely reactive, sentient, and intentional behavior with the latter formalized in terms of inductive inference. This distinction is further studied using simple machine learning benchmarks (navigation in a grid world and the Tower of Hanoi problem) that show how quickly and efficiently adaptive behavior emerges under an inductive form of active inference.
Karl J. Friston, Tommaso Salvatori, Takuya Isomura, Alexander Tschantz, Alex B. Kiefer, Tim Verbelen, Magnus T. Koudahl, Aswin Paul, Thomas Parr, Adeel Razi, Brett J. Kagan, Christopher L. Buckley, Maxwell James D. Ramstead
Neural Comput.6
2024 GenRL: Multimodal-foundation world models for generalization in embodied agents
abstract
Learning generalist embodied agents, able to solve multitudes of tasks in different domains is a long-standing problem. Reinforcement learning (RL) is hard to scale up as it requires a complex reward design for each task. In contrast, language can specify tasks in a more natural way. Current foundation vision-language models (VLMs) generally require fine-tuning or other adaptations to be adopted in embodied contexts, due to the significant domain gap. However, the lack of multimodal data in such domains represents an obstacle to developing foundation models for embodied applications. In this work, we overcome these problems by presenting multimodal-foundation world models, able to connect and align the representation of foundation VLMs with the latent space of generative world models for RL, without any language annotations. The resulting agent learning framework, GenRL, allows one to specify tasks through vision and/or language prompts, ground them in the embodied domain’s dynamics, and learn the corresponding behaviors in imagination. As assessed through large-scale multi-task benchmarking in locomotion and manipulation domains, GenRL enables multi-task generalization from language and visual prompts. Furthermore, by introducing a data-free policy learning strategy, our approach lays the groundwork for foundational policy learning using generative world models. Website, code and data: https://mazpie.github.io/genrl/
Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron C. Courville, Sai Rajeswar
NeurIPS2
2024 Object-Centric Scene Representations Using Active Inference
abstract
Representing a scene and its constituent objects from raw sensory data is a core ability for enabling robots to interact with their environment. In this letter, we propose a novel approach for scene understanding, leveraging an object-centric generative model that enables an agent to infer object category and pose in an allocentric reference frame using active inference, a neuro-inspired framework for action and perception. For evaluating the behavior of an active vision agent, we also propose a new benchmark where, given a target viewpoint of a particular object, the agent needs to find the best matching viewpoint given a workspace with randomly positioned objects in 3D. We demonstrate that our active inference agent is able to balance epistemic foraging and goal-driven behavior, and quantitatively outperforms both supervised and reinforcement learning baselines by more than a factor of two in terms of success rate.
Toon Van de Maele, Tim Verbelen, Pietro Mazzaglia, Stefano Ferraro, Bart Dhoedt
Neural Comput.2
2023 Choreographer: Learning and Adapting Skills in Imagination
Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Alexandre Lacoste, Sai Rajeswar
ICLR2
2023 Mastering the Unsupervised Reinforcement Learning Benchmark from Pixels
abstract
Controlling artificial agents from visual sensory data is an arduous task. Reinforcement learning (RL) algorithms can succeed but require large amounts of interactions between the agent and the environment. To alleviate the issue, unsupervised RL proposes to employ self-supervised interaction and learning, for adapting faster to future tasks. Yet, as shown in the Unsupervised RL Benchmark (URLB; Laskin et al. 2021), whether current unsupervised strategies can improve generalization capabilities is still unclear, especially in visual control settings. In this work, we study the URLB and propose a new method to solve it, using unsupervised model-based RL, for pre-training the agent, and a task-aware fine-tuning strategy combined with a new proposed hybrid planner, Dyna-MPC, to adapt the agent for downstream tasks. On URLB, our method obtains 93.59% overall normalized performance, surpassing previous baselines by a staggering margin. The approach is empirically evaluated through a large-scale empirical study, which we use to validate our design choices and analyze our models. We also show robust performance on the Real-Word RL benchmark, hinting at resiliency to environment perturbations during adaptation. Project website: https://masteringurlb.github.io/
Sai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché, Bart Dhoedt, Aaron C. Courville, Alexandre Lacoste
ICML3
2023 Fusing Event-based Camera and Radar for SLAM Using Spiking Neural Networks with Continual STDP Learning
abstract
This work proposes a first-of-its-kind SLAM architecture fusing an event-based camera and a Frequency Modulated Continuous Wave (FMCW) radar for drone navigation. Each sensor is processed by a bio-inspired Spiking Neural Network (SNN) with continual Spike-Timing-Dependent Plasticity (STDP) learning, as observed in the brain. In contrast to most learning-based SLAM systems, our method does not require any offline training phase, but rather the SNN continuously learns features from the input data on the fly via STDP. At the same time, the SNN outputs are used as feature descriptors for loop closure detection and map correction. We conduct numerous experiments to benchmark our system against state-of-the-art RGB methods and we demonstrate the robustness of our DVS-Radar SLAM approach under strong lighting variations.
Ali Safa, Tim Verbelen, Ilja Ocket, André Bourdoux, Hichem Sahli, Francky Catthoor, Georges Gielen
ICRA2
2022 Curiosity-Driven Exploration via Latent Bayesian Surprise
abstract
The human intrinsic desire to pursue knowledge, also known as curiosity, is considered essential in the process of skill acquisition. With the aid of artificial curiosity, we could equip current techniques for control, such as Reinforcement Learning, with more natural exploration capabilities. A promising approach in this respect has consisted of using Bayesian surprise on model parameters, i.e. a metric for the difference between prior and posterior beliefs, to favour exploration. In this contribution, we propose to apply Bayesian surprise in a latent space representing the agent’s current understanding of the dynamics of the system, drastically reducing the computational costs. We extensively evaluate our method by measuring the agent's performance in terms of environment exploration, for continuous tasks, and looking at the game scores achieved, for video games. Our model is computationally cheap and compares positively with current state-of-the-art methods on several problems. We also investigate the effects caused by stochasticity in the environment, which is often a failure case for curiosity-driven agents. In this regime, the results suggest that our approach is resilient to stochastic transitions.
Pietro Mazzaglia, Ozan Çatal, Tim Verbelen, Bart Dhoedt
AAAI3
2022 Iterative neural networks for adaptive inference on resource-constrained devices
Sam Leroux, Tim Verbelen, Pieter Simoens, Bart Dhoedt
Neural Comput. Appl.2
2021 LatentSLAM: unsupervised multi-sensor representation learning for localization and mapping
abstract
Biologically inspired algorithms for simultaneous localization and mapping (SLAM) such as RatSLAM have been shown to yield effective and robust robot navigation in both indoor and outdoor environments. One drawback however is the sensitivity to perceptual aliasing due to the template matching of low-dimensional sensory templates. In this paper, we propose an unsupervised representation learning method that yields low-dimensional latent state descriptors that can be used for RatSLAM. Our method is sensor agnostic and can be applied to any sensor modality, as we illustrate for camera images, radar range-doppler maps and lidar scans. We also show how combining multiple sensors can increase the robustness, by reducing the number of false matches. We evaluate on a dataset captured with a mobile robot navigating in a warehouse-like environment, moving through different aisles with similar appearance, making it hard for the SLAM algorithms to disambiguate locations.
Ozan Çatal, Wouter Jansen, Tim Verbelen, Bart Dhoedt, Jan Steckel
ICRA3
2021 Dynamic Narrowing of VAE Bottlenecks Using GECO and L0 Regularization
abstract
When designing variational autoencoders (VAEs) or other types of latent space models, the dimensionality of the latent space is typically defined upfront. In this process, it is possible that the number of dimensions is under- or overprovisioned for the application at hand. In case the dimensionality is not predefined, this parameter is usually determined using time- and resource-consuming cross-validation. For these reasons we have developed a technique to shrink the latent space dimensionality of VAEs automatically and on-the-fty during training using Generalized ELBO with Constrained Optimization (GECO) and the$L_{0}$-Augment-REINFORcE-Merge ($L_{0}$-ARM) gradient estimator. The GECO optimizer ensures that we are not violating a predefined upper bound on the reconstruction error. This paper presents the algorithmic details of our method along with experimental results on five different datasets. We find that our training procedure is stable and that the latent space can be pruned effectively without violating the GECO constraints.
Cedric De Boom, Samuel Wauthier, Tim Verbelen, Bart Dhoedt
IJCNN3
2021 Contrastive Active Inference
abstract
Active inference is a unifying theory for perception and action resting upon the idea that the brain maintains an internal model of the world by minimizing free energy. From a behavioral perspective, active inference agents can be seen as self-evidencing beings that act to fulfill their optimistic predictions, namely preferred outcomes or goals. In contrast, reinforcement learning requires human-designed rewards to accomplish any desired outcome. Although active inference could provide a more natural self-supervised objective for control, its applicability has been limited because of the shortcomings in scaling the approach to complex environments. In this work, we propose a contrastive objective for active inference that strongly reduces the computational burden in learning the agent's generative model and planning future actions. Our method performs notably better than likelihood-based active inference in image-based tasks, while also being computationally cheaper and easier to train. We compare to reinforcement learning agents that have access to human-designed reward functions, showing that our approach closely matches their performance. Finally, we also show that contrastive methods perform significantly better in the case of distractors in the environment and that our method is able to generalize goals to variations in the background.
Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt
NeurIPS2
2021 Leveraging the Bhattacharyya coefficient for uncertainty quantification in deep neural networks
abstract
Abstract Modern deep learning models achieve state-of-the-art results for many tasks in computer vision, such as image classification and segmentation. However, its adoption into high-risk applications, e.g. automated medical diagnosis systems, happens at a slow pace. One of the main reasons for this is that regular neural networks do not capture uncertainty. To assess uncertainty in classification, several techniques have been proposed casting neural network approaches in a Bayesian setting. Amongst these techniques, Monte Carlo dropout is by far the most popular. This particular technique estimates the moments of the output distribution through sampling with different dropout masks. The output uncertainty of a neural network is then approximated as the sample variance. In this paper, we highlight the limitations of such a variance-based uncertainty metric and propose an novel approach. Our approach is based on the overlap between output distributions of different classes. We show that our technique leads to a better approximation of the inter-class output confusion. We illustrate the advantages of our method using benchmark datasets. In addition, we apply our metric to skin lesion classification—a real-world use case—and show that this yields promising results.
Pieter Van Molle, Tim Verbelen, Bert Vankeirsbilck, Jonas De Vylder, Bart Diricx, Tom Kimpe, Pieter Simoens, Bart Dhoedt
Neural Comput. Appl.2
2021 Robot navigation as hierarchical active inference
Ozan Çatal, Tim Verbelen, Toon Van de Maele, Bart Dhoedt, Adam Safron
Neural Networks2
2020 Learning Perception and Planning With Deep Active Inference
abstract
Active inference is a process theory of the brain that states that all living organisms infer actions in order to minimize their (expected) free energy. However, current experiments are limited to predefined, often discrete, state spaces. In this paper we use recent advances in deep learning to learn the state space and approximate the necessary probability distributions to engage in active inference.
Ozan Çatal, Tim Verbelen, Johannes Nauta, Cedric De Boom, Bart Dhoedt
ICASSP2
2020 Anomaly Detection for Autonomous Guided Vehicles using Bayesian Surprise
abstract
As warehouses, storage facilities and factories become more expanded and equipped with smart devices, there is a substantial need for rapid, intelligent and autonomous detection of unusual and potentially hazardous situations, also called anomalies. In particular for Autonomous Guided Vehicles (AGVs) that drive around these premises independently, unforeseen obstructions along their path-e.g. a cardboard box in the middle of a corridor or bumps in the floor-and sudden or unexpected actions executed by personnel-e.g. someone walking in a restricted area-make it hard for AGVs to navigate safely. We therefore propose a novel approach to detect such anomalies in an unsupervised manner by measuring Bayesian surprise: whenever an event is observed that does not align with the agent's prior knowledge of the world, this event is deemed surprising and could indicate an anomaly. This paper lays out the details on how to learn both the prior and posterior models of an AGV that drives around a warehouse and observes the environment through an RGBD camera. In the experiments we show that our Bayesian surprise approach outperforms a baseline that is traditionally used to detect anomalies in sequences of images.
Ozan Çatal, Sam Leroux, Cedric De Boom, Tim Verbelen, Bart Dhoedt
IROS4
2020 Training binary neural networks with knowledge transfer
Sam Leroux, Bert Vankeirsbilck, Tim Verbelen, Pieter Simoens, Bart Dhoedt
Neurocomputing3
2019 Learning to Grasp Arbitrary Household Objects from a Single Demonstration
abstract
Upon the advent of Industry 4.0, collaborative robotics and intelligent automation gain more and more traction for enterprises to improve their production processes. In order to adapt to this trend, new programming, learning and collaborative techniques are investigated. Program-bydemonstration is one of the techniques that aim to reduce the burden of manually programming collaborative robots. However, this is often limited to teaching to grasp at a certain position, rather than grasping a certain object. In this paper, we propose a method that learns to grasp an arbitrary object from visual input. While other learning-based approaches for robotic grasping require collecting a large dataset, manually or automatically labeled in a real or simulated world, our approach requires a single demonstration. We present results on grasping various objects with the Franka Panda collaborative robot after capturing a single image from a wrist mounted RGB camera. From this image we learn a robot controller with a convolutional neural network to adapt to changes in the object's position and rotation with less than 5 minutes of training time on a NVIDIA Titan X GPU, achieving over 90% grasp success rate.
Elias De Coninck, Tim Verbelen, Pieter Van Molle, Pieter Simoens, Bart Dhoedt
IROS2
2019 Multi-fidelity deep neural networks for adaptive inference in the internet of multimedia things
Sam Leroux, Steven Bohez, Elias De Coninck, Pieter Van Molle, Bert Vankeirsbilck, Tim Verbelen, Pieter Simoens, Bart Dhoedt
Future Gener. Comput. Syst.6
2018 DIANNE: a modular framework for designing, training and deploying deep neural networks on heterogeneous distributed infrastructure
Elias De Coninck, Steven Bohez, Sam Leroux, Tim Verbelen, Bert Vankeirsbilck, Pieter Simoens, Bart Dhoedt
J. Syst. Softw.4
2017 Sensor fusion for robot control through deep reinforcement learning
abstract
Deep reinforcement learning is becoming increasingly popular for robot control algorithms, with the aim for a robot to self-learn useful feature representations from unstructured sensory input leading to the optimal actuation policy. In addition to sensors mounted on the robot, sensors might also be deployed in the environment, although these might need to be accessed via an unreliable wireless connection. In this paper, we demonstrate deep neural network architectures that are able to fuse information generated by multiple sensors and are robust to sensor failures at runtime. We evaluate our method on a search and pick task for a robot both in simulation and the real world.
Steven Bohez, Tim Verbelen, Elias De Coninck, Bert Vankeirsbilck, Pieter Simoens, Bart Dhoedt
IROS2
2017 The cascading neural network: building the Internet of Smart Things
Sam Leroux, Steven Bohez, Elias De Coninck, Tim Verbelen, Bert Vankeirsbilck, Pieter Simoens, Bart Dhoedt
Knowl. Inf. Syst.4
2016 Multi-fidelity matryoshka neural networks for constrained IoT devices
abstract
Using deep neural networks on resource constrained devices is a trending topic in neural network research. Various techniques for compressing neural networks have been proposed that allow evaluating a large neural network on a device with limited memory and processing power. These approaches usually generate a single compressed student network based on a larger teacher network. In some cases a more dynamic trade-off may be desired. In this paper we trained a sequence of increasingly large networks where each network is constrained to contain the unmodified features of all smaller networks. The weight matrix of the largest network has submatrices that correspond to the weight matrices of each of the smaller networks. This technique allows us to keep the parameters of several networks in memory while having the same memory footprint as the single largest network. A trade-off between accuracy and speed can be made at runtime. The proposed approach is validated on two image classification tasks running on a real-world Internet-of-Things (IoT) device.
Sam Leroux, Steven Bohez, Elias De Coninck, Tim Verbelen, Bert Vankeirsbilck, Pieter Simoens, Bart Dhoedt
IJCNN4
2016 Mobile device power models for energy efficient dynamic offloading at runtime
Farhan Azmat Ali, Pieter Simoens, Tim Verbelen, Piet Demeester, Bart Dhoedt
J. Syst. Softw.3
2016 Dynamic auto-scaling and scheduling of deadline constrained service workloads on IaaS clouds
Elias De Coninck, Tim Verbelen, Bert Vankeirsbilck, Steven Bohez, Pieter Simoens, Bart Dhoedt
J. Syst. Softw.2
2015 Resource-constrained classification using a cascade of neural network layers
abstract
Deep neural networks are the state of the art technique for a wide variety of classification problems. Although deeper networks are able to make more accurate classifications, the value brought by an additional hidden layer diminishes rapidly. Even shallow networks are able to achieve relatively good results on various classification problems. Only for a small subset of the samples do the deeper layers make a significant difference. We describe an architecture in which only the samples that can not be classified with a sufficient confidence by a shallow network have to be processed by the deeper layers. Instead of training a network with one output layer at the end of the network, we train several output layers, one for each hidden layer. When an output layer is sufficiently confident in this result, we stop propagating at this layer and the deeper layers need not be evaluated. The choice of a threshold confidence value allows us to trade-off accuracy and speed.
Sam Leroux, Steven Bohez, Tim Verbelen, Bert Vankeirsbilck, Pieter Simoens, Bart Dhoedt
IJCNN3
2014 Management of crowdsourced first-person video: street view live
abstract
We present a framework for large-scale crowdsourcing of first-person viewpoint videos recorded on mobile devices. Collecting videos at a massive scale poses a number of major issues in terms of network planning. To improve the scalability with regards to the number of users, videos and geographical area and better cope with restrictions on storage, bandwidth and processing power, the framework is distributed and based on the two-layer cloudlet architecture. To mitigate the limited bandwidth in the access network, a set of decision algorithms is constructed and evaluated that are able to filter out irrelevant videos based on their metadata and given selection criteria. To illustrate the crowdsourcing framework, we present Street View Live, an application for presenting videos based on location, similar to the popular Google Street View but with up-to-date videos covering the location instead of possibly outdated images. In order to have an up-to-date view of every location, the video collection is continuously extended and updated by crowdsourcing videos from mobile devices.
Steven Bohez, Jens Mostaert, Tim Verbelen, Pieter Simoens, Bart Dhoedt
MUM3
2014 Adaptive deployment and configuration for mobile augmented reality in the cloudlet
Tim Verbelen, Pieter Simoens, Filip De Turck, Bart Dhoedt
J. Netw. Comput. Appl.1
2013 Quality of experience driven control of interactive media stream parameters
Bert Vankeirsbilck, Tim Verbelen, Dieter Verslype, Nicolas Staelens, Filip De Turck, Piet Demeester, Bart Dhoedt
IM2
2013 Graph partitioning algorithms for optimizing software deployment in mobile cloud computing
Tim Verbelen, Tim Stevens, Filip De Turck, Bart Dhoedt
Future Gener. Comput. Syst.1
2012 A component-based approach towards mobile distributed and collaborative PTAM
abstract
Having numerous sensors on-board, smartphones have rapidly become a very attractive platform for augmented reality applications. Although the computational resources of mobile devices grow, they still cannot match commonly available desktop hardware, which results in downscaled versions of well known computer vision techniques that sacrifice accuracy for speed. We propose a component-based approach towards mobile augmented reality applications, where components can be configured and distributed at runtime, resulting in a performance increase by offloading CPU intensive tasks to a server in the network. By sharing distributed components between multiple users, collaborative AR applications can easily be developed. In this poster, we present a component-based implementation of the Parallel Tracking And Mapping (PTAM) algorithm, enabling to distribute components to achieve a mobile, distributed version of the original PTAM algorithm, as well as a collaborative scenario.
Tim Verbelen, Pieter Simoens, Filip De Turck, Bart Dhoedt
ISMAR1
2012 AIOLOS: Middleware for improving mobile application performance through cyber foraging
Tim Verbelen, Pieter Simoens, Filip De Turck, Bart Dhoedt
J. Syst. Softw.1
2011 Dynamic deployment and quality adaptation for mobile augmented reality applications
Tim Verbelen, Tim Stevens, Pieter Simoens, Filip De Turck, Bart Dhoedt
J. Syst. Softw.1