EDBT 2026 Demo / reviewers in the wild / expert
Markus Grotz
dblp:173/7849
· DBLP profile ↗
10ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0001-7257-5872ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic ManipulationabstractRobotic manipulation systems operating in diverse, dynamic environments must exhibit three critical abilities: multitask interaction, generalization to unseen scenarios, and spatial memory. While significant progress has been made in robotic manipulation, existing approaches often fall short in generalization to complex environmental variations and addressing memory-dependent tasks. To bridge this gap, we introduce **SAM2Act**, a multi-view robotic transformer-based policy that leverages multi-resolution upsampling with visual representations from large-scale foundation model. SAM2Act achieves a state-of-the-art average success rate of **86.8% across 18 tasks** in the RLBench benchmark, and demonstrates robust generalization on *The Colosseum* benchmark, with only a **4.3% performance gap** under diverse environmental perturbations. Building on this foundation, we propose **SAM2Act+**, a memory-based architecture inspired by SAM2, which incorporates a memory bank, an encoder, and an attention mechanism to enhance spatial memory. To address the need for evaluating memory-dependent tasks, we introduce ***MemoryBench***, a novel benchmark designed to assess spatial memory and action recall in robotic manipulation. SAM2Act+ achieves an average success rate of **94.3% on memory-based tasks** in *MemoryBench*, significantly outperforming existing approaches and pushing the boundaries of memory-based robotic systems.
Project page: [sam2act.github.io](https://sam2act.github.io/). Haoquan Fang, Markus Grotz, Wilbert Pumacay, Yi Ru Wang, Dieter Fox, Ranjay Krishna, Jiafei Duan |
ICML | 2 |
| 2025 | TWIN: Two-handed Intelligent Benchmark for Bimanual ManipulationabstractBimanual manipulation is challenging due to precise spatial and temporal coordination required between two arms. While there exist several real-world bimanual systems, there is a lack of simulated benchmarks with a large task diversity for systematically studying bimanual capabilities across a wide range of tabletop tasks. This paper addresses the gap by presenting a benchmark for bimanual manipulation. A key functionality is the ability to autonomously generate training data without the necessity of human demonstrations to the robot. We open-source our code and benchmark, which comprises 13 new tasks with 23 unique task variations, each requiring a high degree of coordination and adaptability. To initiate the benchmark, we extended multiple state-of-the-art techniques to the domain of bimanual manipulation. The project website with code is available at: http://bimanual.github.io. Markus Grotz, Mohit Shridhar, Yu-Wei Chao, Tamim Asfour, Dieter Fox |
ICRA | 1 |
| 2025 | OptiGrasp: Optimized Grasp Pose Detection Using RGB Images for Warehouse Picking RobotsabstractIn warehouse environments, robots require robust picking capabilities to manage a wide variety of objects. Effective deployment demands minimal hardware, strong generalization to new products, and resilience in diverse settings. Current methods often rely on depth sensors for structural information, which suffer from high costs, complex setups, and technical limitations. Inspired by recent advancements in computer vision, we propose an innovative approach that leverages foundation models to enhance suction grasping using only RGB images. Trained solely on a synthetic dataset, our method generalizes its grasp prediction capabilities to real-world robots and a diverse range of novel objects not included in the training set. Our network achieves an 81.9% success rate in real-world applications. The project website with code and data will be available at http://optigrasp.github.io. Soofiyan Atar, Yi Li 0038, Markus Grotz, Michael Wolf, Dieter Fox, Joshua R. Smith 0001 |
IROS | 3 |
| 2025 | TetraGrip: Sensor-Driven Multi-Suction Reactive Object Manipulation in Cluttered ScenesabstractWarehouse robotic systems equipped with vacuum grippers must reliably grasp a diverse range of objects from densely packed shelves. However, these environments present significant challenges, including occlusions, diverse object orientations, stacked and obstructed items, and surfaces that are difficult to suction. We introduce TetraGrip, a novel vacuum-based grasping strategy featuring four suction cups mounted on linear actuators. Each actuator is equipped with an optical time-of-flight (ToF) proximity sensor, enabling reactive grasping.We evaluate TetraGrip in a warehouse-style setting, demonstrating its ability to manipulate objects in stacked and obstructed configurations. Our results show that our RL-based policy improves picking success in stacked-object scenarios by 22.9% compared to a single-suction gripper. Additionally, we demonstrate that TetraGrip can successfully grasp objects in scenarios where a single-suction gripper fails due to physical limitations, specifically in two cases: (1) picking an object occluded by another object and (2) retrieving an object in a complex scenario. These findings highlight the advantages of multi-actuated, suction-based grasping in unstructured warehouse environments. The project website is available at: https://tetragrip.github.io/. Paolo Torrado, Joshua Levin, Markus Grotz, Joshua R. Smith 0001 |
IROS | 3 |
| 2022 | Learning Symbolic Failure Detection for Grasping and Mobile Manipulation TasksabstractThe ability to detect failure during task execution and to recover from failure is vital for autonomous robots performing tasks in previously unknown environments. In this paper, we present an approach for failure detection during the execution of grasping and mobile manipulation tasks by a humanoid robot. The approach combines multi-modal sensory information consisting of proprioceptive, force and visual information to learn task models from multiple successful task executions, in order to detect failures and to externalize them for humans in an interpretable way. To this end, we define symbolic action predicates based on multi-modal sensory information to allow high-level state estimation based on action-specific decision trees. To allow symbolic failure detection, we then learn task models that are represented as Markov chains. We evaluated the approach in several pick-and-place and mobile manipulation tasks performed by a humanoid robot in a decommissioning and a household scenario. The evaluation shows that the learned task models are capable of detecting failure with an F1-score of 93 %. Patrick Hegemann, Tim Zechmeister, Markus Grotz, Kevin Hitzler, Tamim Asfour |
IROS | 3 |
| 2022 | BlueSky: Combining Task Planning and Activity-Centric Access Control for Assistive Humanoid RobotsabstractIn the not too distant future, assistive humanoid robots will provide versatile assistance for coping with everyday life. In their interactions with humans, not only safety, but also security and privacy issues need to be considered. In this Blue Sky paper, we therefore argue that it is time to bring task planning and execution as a well-established field of robotics with access and usage control in the field of security and privacy closer together. In particular, the recently proposed activity-based view on access and usage control provides a promising approach to bridge the gap between these two perspectives. We argue that humanoid robots provide for specific challenges due to their task-universality and their use in both, private and public spaces. Furthermore, they are socially connected to various parties and require policy creation at runtime due to learning. We contribute first attempts on the architecture and enforcement layer as well as on joint modeling, and discuss challenges and a research roadmap also for the policy and objectives layer. We conclude that the underlying combination of decentralized systems' and smart environments' research aspects provides for a rich source of challenges that need to be addressed on the road to deployment. Saskia Bayreuther, Florian Jacob, Markus Grotz, Rainer Kartmann, Fabian Tërnava, Fabian Paus, Hannes Hartenstein, Tamim Asfour |
SACMAT | 3 |
| 2021 | Vision-Based Robotic Pushing and Grasping for Stone Sample Collection under Computing Resource ConstraintsabstractIncreasing the robustness of grasping actions and the recovery from failure is key to improving a robot’s autonomy. Endowing robots with the ability to robustly grasp and manipulate unknown difficult objects such as stones is required for sample collection in unknown environments. In this paper, we present a complete system for robust grasping of stones, which integrates stone segmentation based on depth information, the generation of grasp hypotheses and pushing actions as well as their execution. In particular, our system has been designed to solve these tasks on robots with limited computing resources. We evaluate the performance in real robot experiments in the context of stone sample collection. The results show that such a challenging task is achievable under computing resource constraints. Raphael Grimm, Markus Grotz, Simon Ottenhaus, Tamim Asfour |
ICRA | 2 |
| 2017 | Autonomous view selection and gaze stabilization for humanoid robotsabstractTo increase the autonomy of humanoid robots, the visual perception must support the efficient collection and interpretation of visual scene cues by providing task-dependent information. Active vision systems allow to extend the observable workspace by employing active gaze control, i.e. by shifting the gaze to relevant areas in the scene. When moving the eyes, stabilization of the camera images is crucial for successful task execution. In this paper, we present an active vision system for task-oriented selection of view directions and gaze stabilization to enable a humanoid robot to robustly perform vision-based tasks. We investigate the interaction between a gaze stabilization controller and view planning to select the next best view direction based on saliency maps which encode task-relevant information. We demonstrate the performance of the systems in a real world scenario, in which a humanoid robot is performing vision-based grasping while moving, a task that would not be possible without the combination of view selection and gaze stabilization. Markus Grotz, Timothee Habra, Renaud Ronsse, Tamim Asfour |
IROS | 1 |
| 2016 | Towards a hierarchy of loco-manipulation affordancesabstractWe propose a formalism for the hierarchical representation of affordances. Starting with a perceived model of the environment consisting of geometric primitives like planes or cylinders, we define a hierarchical system for affordance extraction whose foundation are elementary power grasp affordances. Higher-level affordances, e.g. bimanual affordances, result from combining lower-level affordances with additional properties concerning the underlying geometric primitives of the scene. We model affordances as continuous certainty functions taking into account properties of the environmental elements and the perceiving robot's embodiment. The developed formalism is regarded as the basis for the description of whole-body affordances, i.e. affordances associated with whole-body actions. The proposed formalism was implemented and experimentally evaluated in multiple scenarios based on RGB-D camera data. The feasibility of the approach is demonstrated on a real robotic platform. Peter Kaiser 0001, Eren Erdal Aksoy, Markus Grotz, Tamim Asfour |
IROS | 3 |
| 2015 | On the Dualities Between Grasping and Whole-Body Loco-Manipulation Tasks
Tamim Asfour, Júlia Borràs Sol, Christian Mandery, Peter Kaiser 0001, Eren Erdal Aksoy, Markus Grotz |
ISRR (2) | 6 |