Fabian Otto

dblp:284/0547 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0003-3484-1054ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 46% Robot manipulation · 26% Motion planning and robot control · 20%
Network and information security
1 paper
Privacy and data protection · 91% Usable security · 9%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot manipulation
grasping
0.912025
Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning · ICRA 2025
Machine learning › Reinforcement learning
imitation learning
0.912025
BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning · NeurIPS 2025
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.912025
Efficient Off-Policy Learning for High-Dimensional Action Spaces · ICLR 2025
Robotics › Robot manipulation › grasping › pre-grasp manipulation
pre-grasp planning
0.912025
Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning · ICRA 2025
Robotics › Motion planning and robot control
robot control
0.912025
BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning · NeurIPS 2025
Robotics › Motion planning and robot control
trajectory planning
0.912025
BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning · NeurIPS 2025
Machine learning › Reinforcement learning
episodic reinforcement learning
0.812024
Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning · ICLR 2024
Machine learning › Reinforcement learning
exploration
0.812024
Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning · ICLR 2024
Machine learning › Reinforcement learning › exploration
parameter space exploration
0.812024
Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning · ICLR 2024
Privacy and data protection › surveillance
client-side scanning
0.712023
Attitudes towards Client-Side Scanning for CSAM, Terrorism, Drug Trafficking, Drug Use and Tax Evasion in Germany · SP 2023
Privacy and data protection
privacy perceptions
0.712023
Attitudes towards Client-Side Scanning for CSAM, Terrorism, Drug Trafficking, Drug Use and Tax Evasion in Germany · SP 2023
Privacy and data protection
surveillance
0.712023
Attitudes towards Client-Side Scanning for CSAM, Terrorism, Drug Trafficking, Drug Use and Tax Evasion in Germany · SP 2023
Machine learning › Reinforcement learning › policy optimization
trust region methods
0.512021
Differentiable Trust Region Layers for Deep Reinforcement Learning · ICLR 2021
Computer vision › 3D vision
3d scene reconstruction
0.312025
Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning · ICRA 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › heuristic search
heuristic search planning
0.312025
Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning · ICRA 2025
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
importance sampling
0.312025
Efficient Off-Policy Learning for High-Dimensional Action Spaces · ICLR 2025
Robotics › Motion planning and robot control
motion planning
0.312025
Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning · ICRA 2025

Methods — techniques the papers use, named apart from their topics

weighted importance sampling · 0.9vector quantization · 0.9twin value function networks · 0.9parallel decoding · 0.9image segmentation · 0.9heuristic-based pushing · 0.9b-spline encoding · 0.93d reconstruction · 0.9step-based policy update · 0.8survey · 0.7
YearPublicationVenuePosition
2025 Efficient Off-Policy Learning for High-Dimensional Action Spaces
abstract
Existing off-policy reinforcement learning algorithms often rely on an explicit state-action-value function representation, which can be problematic in high-dimensional action spaces due to the curse of dimensionality. This reliance results in data inefficiency as maintaining a state-action-value function in such spaces is challenging. We present an efficient approach that utilizes only a state-value function as the critic for off-policy deep reinforcement learning. This approach, which we refer to as Vlearn, effectively circumvents the limitations of existing methods by eliminating the necessity for an explicit state-action-value function. To this end, we leverage a weighted importance sampling loss for learning deep value functions from off-policy data. While this is common for linear methods, it has not been combined with deep value function networks. This transfer to deep methods is not straightforward and requires novel design choices such as robust policy updates, twin value function networks to avoid an optimization bias, and importance weight clipping. We also present a novel analysis of the variance of our estimate compared to commonly used importance sampling estimators such as V-trace. Our approach improves sample complexity as well as final performance and ensures consistent and robust performance across various benchmark tasks. Eliminating the state-action-value function in Vlearn facilitates a streamlined learning process, yielding high-return agents.
Fabian Otto, Philipp Becker, Ngo Anh Vien, Gerhard Neumann
ICLR1
2025 Robust Optical Transceiver Manipulation in Cluttered Cable Environments Using 3D Scene Understanding and Planning
abstract
Robotic manipulation in cluttered environments presents significant challenges, particularly when the clutter includes thin, deformable objects like cables, which complicate perception and decision-making processes. In the context of datacenters, the automation of networking tasks often involves the manipulation of optical transceivers within densely packed cable configurations. Such environments are characterized by an abundance of delicate, overlapping, and intersecting cables, leading to frequent occlusions. This paper introduces an innovative system designed for the manipulation of optical transceivers in environments cluttered by cables. Our integrated approach combines advanced 3D scene understanding with a heuristic-based pushing policy to effectively manipulate optical transceivers amidst clutter. The system's perception component utilizes image segmentation and 3D reconstruction to accurately model the transceivers and surrounding cables. Meanwhile, the planning aspect employs a search algorithm with task-specific heuristics, to navigate the gripper, displace obstructing cables, and safely achieve a precise pre-grasp position in front of the target transceiver. We have conducted extensive evaluations of our methodology in both simulated and real-world settings, demonstrating its high success rates, robustness, and proficiency in addressing the unique challenges posed by cable-occluded environments within datacenters.
Iason Sarantopoulos, Bohong Weng, Sicheng Xu, Jiaolong Yang, Xin Tong 0001, Fabian Otto, David Sweeney, Andromachi Chatzieleftheriou, Antony I. T. Rowstron
ICRA8
2025 BEAST: Efficient Tokenization of B-Splines Encoded Action Sequences for Imitation Learning
abstract
We present the B-spline Encoded Action Sequence Tokenizer (BEAST), a novel action tokenizer that encodes action sequences into compact discrete or continuous tokens using B-splines. In contrast to existing action tokenizers based on vector quantization or byte pair encoding, BEAST requires no separate tokenizer training and consistently produces tokens of uniform length, enabling fast action sequence generation via parallel decoding. Leveraging our B-spline formulation, BEAST inherently ensures generating smooth trajectories without discontinuities between adjacent segments. We extensively evaluate BEAST by integrating it with three distinct model architectures: a Variational Autoencoder (VAE) with continuous tokens, a decoder-only Transformer with discrete tokens, and Florence-2, a pretrained Vision-Language Model with an encoder-decoder architecture, demonstrating BEAST's compatibility and scalability with large pretrained models. We evaluate BEAST across three established benchmarks consisting of 166 simulated tasks and on three distinct robot settings with a total of 8 real-world tasks. Experimental results demonstrate that BEAST (i) significantly reduces both training and inference computational costs, and (ii) consistently generates smooth, high-frequency control signals suitable for continuous control tasks while (iii) reliably achieves competitive task success rates compared to state-of-the-art methods.
Hongyi Zhou, Weiran Liao, Xi Huang 0005, Yucheng Tang, Fabian Otto, Xiaogang Jia, Xinkai Jiang, Simon Hilber, Ömer Erdinç Yagmurlu, Nils Blank, Moritz Reuss, Rudolf Lioutikov
NeurIPS5
2024 Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning
abstract
Current advancements in reinforcement learning (RL) have predominantly focused on learning step-based policies that generate actions for each perceived state. While these methods efficiently leverage step information from environmental interaction, they often ignore the temporal correlation between actions, resulting in inefficient exploration and unsmooth trajectories that are challenging to implement on real hardware. Episodic RL (ERL) seeks to overcome these challenges by exploring in parameters space that capture the correlation of actions. However, these approaches typically compromise data efficiency, as they treat trajectories as opaque black boxes. In this work, we introduce a novel ERL algorithm, Temporally-Correlated Episodic RL (TCE), which effectively utilizes step information in episodic policy updates, opening the 'black box' in existing ERL methods while retaining the smooth and consistent exploration in parameter space. TCE synergistically combines the advantages of step-based and episodic RL, achieving comparable performance to recent ERL methods while maintaining data efficiency akin to state-of-the-art (SoTA) step-based RL.
Hongyi Zhou, Dominik Roth, Serge Thilges, Fabian Otto, Rudolf Lioutikov, Gerhard Neumann
ICLR5
2023 Attitudes towards Client-Side Scanning for CSAM, Terrorism, Drug Trafficking, Drug Use and Tax Evasion in Germany
abstract
In recent years, there have been a rising number of legislative efforts and proposed technical measures to weaken privacy-preserving technology, with the stated goal of countering serious crimes like child abuse. One of these proposed measures is Client-Side Scanning (CSS). CSS has been hotly debated both in the context of Apple stating their intention to deploy it in 2021 as well as EU legislation being proposed in 2022. Both sides of the argument state that they are working in the best interests of the people. To shed some light on this, we conducted a survey with a representative sample of German citizens. We investigated the general acceptance of CSS vs cloud-based scanning for different types of crimes and analyzed how trust in the German government and companies such as Google and Apple influenced our participants’ views. We found that, by and large, the majority of participants were willing to accept CSS measures to combat serious crimes such as child abuse or terrorism, but support dropped significantly for other illegal activities. However, the majority of participants who supported CSS were also worried about potential abuse, with only 20% stating that they were not concerned. These results suggest that many of our participants would be willing to have their devices scanned and accept some risks in the hope of aiding law enforcement. In our analysis, we argue that there are good reasons to not see this as a carte blanche for the introduction of CSS but as a call to action for the S&P community. More research is needed into how a population’s desire to prevent serious crime online can be achieved while mitigating the risks to privacy and society.
Lisa Geierhaas, Fabian Otto, Maximilian Häring, Matthew Smith 0001
SP2
2021 Differentiable Trust Region Layers for Deep Reinforcement Learning
Fabian Otto, Philipp Becker, Ngo Anh Vien, Hanna Carolin Ziesche, Gerhard Neumann
ICLR1