EDBT 2026 Demo / reviewers in the wild / expert
Jeremie Papon
dblp:26/11103
· DBLP profile ↗
14ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-authorArtificial intelligence and machine learning · 8 · 3 first-authorSystems, architecture and hardware · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 53% Segmentation and scene understanding · 44% Robot manipulation · 3% | |
| Computer graphics and multimedia
2 papers |
Image and video processing · 54% Geometric modeling and processing · 46% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d scene understanding |
0.4 | 2 | 2015 | Semantic Pose Using Deep Networks Trained on Synthetic RGB-D · ICCV 2015 Voxel Cloud Connectivity Segmentation - Supervoxels for Point Clouds · CVPR 2013 |
Computer vision › 3D vision
point cloud segmentation |
0.4 | 2 | 2015 | Constrained planar cuts - Object partitioning for point clouds · CVPR 2015 Voxel Cloud Connectivity Segmentation - Supervoxels for Point Clouds · CVPR 2013 |
Computer vision › Segmentation and scene understanding › scene understanding
indoor scene understanding |
0.2 | 1 | 2015 | Semantic Pose Using Deep Networks Trained on Synthetic RGB-D · ICCV 2015 |
Computer vision › 3D vision
object pose estimation |
0.2 | 1 | 2015 | Semantic Pose Using Deep Networks Trained on Synthetic RGB-D · ICCV 2015 |
Image and video processing › image segmentation
shape segmentation |
0.2 | 1 | 2015 | Constrained planar cuts - Object partitioning for point clouds · CVPR 2015 |
Computer vision › Segmentation and scene understanding
3d point cloud segmentation |
0.2 | 1 | 2014 | Convexity based object partitioning for robot applications · ICRA 2014 |
Computer vision › Segmentation and scene understanding
part segmentation |
0.2 | 1 | 2014 | Convexity based object partitioning for robot applications · ICRA 2014 |
Geometric modeling and processing › point cloud processing
point cloud segmentation |
0.2 | 1 | 2014 | Object Partitioning Using Local Convexity · CVPR 2014 |
Computer vision › Segmentation and scene understanding › 3d segmentation
supervoxel segmentation |
0.2 | 1 | 2013 | Voxel Cloud Connectivity Segmentation - Supervoxels for Point Clouds · CVPR 2013 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.1 | 1 | 2015 | Semantic Pose Using Deep Networks Trained on Synthetic RGB-D · ICCV 2015 |
Robotics › Robot manipulation
grasping |
0.1 | 1 | 2014 | Convexity based object partitioning for robot applications · ICRA 2014 |
Methods — techniques the papers use, named apart from their topics
local concavity graph · 0.4greedy graph cut · 0.4transfer learning · 0.2synthetic data rendering · 0.2convolutional neural network · 0.2voxel grid · 0.2region growing · 0.2convexity analysis · 0.2convex-concave edge classification · 0.2adjacency graph of surface patches · 0.2adjacency graph · 0.2voxel cloud connectivity · 0.2over-segmentation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Teaching a Robot the Semantics of Assembly TasksabstractWe present a three-level cognitive system in a learning by demonstration context. The system allows for learning and transfer on the sensorimotor level as well as the planning level. The fundamentally different data structures associated with these two levels are connected by an efficient mid-level representation based on so-called “semantic event chains.” We describe details of the representations and quantify the effect of the associated learning procedures for each level under different amounts of noise. Moreover, we demonstrate the performance of the overall system by three demonstrations that have been performed at a project review. The described system has a technical readiness level (TRL) of 4, which in an ongoing follow-up project will be raised to TRL 6. Thiusius Rajeeth Savarimuthu, Anders Glent Buch, Christian Schlette, Nils Wantia, Jürgen Roßmann, David Martínez Martínez, Guillem Alenyà, Carme Torras, Ales Ude, Bojan Nemec, Aljaz Kramberger, Florentin Wörgötter, Eren Erdal Aksoy, Jeremie Papon, Simon Haller, Justus H. Piater, Norbert Krüger |
IEEE Trans. Syst. Man Cybern. Syst. | 14 |
| 2017 | Task-oriented grasping with semantic and geometric scene understandingabstractWe present a task-oriented grasp model, that encodes grasps that are configurationally compatible with a given task. For instance, if the task is to pour liquid from a container, the model encodes grasps that leave the opening of the container unobstructed. The model consists of two independent agents: First, a geometric grasp model that computes, from a depth image, a distribution of 6D grasp poses for which the shape of the gripper matches the shape of the underlying surface. The model relies on a dictionary of geometric object parts annotated with workable gripper poses and preshape parameters. It is learned from experience via kinesthetic teaching. The second agent is a CNN-based semantic model that identifies grasp-suitable regions in a depth image: regions where a grasp will not impede the execution of the task. The semantic model allows us to encode relationships such as “grasp from the handle.” A key element of this work is to use a deep network to integrate contextual task cues, and defer the structured-output problem of gripper pose computation to an explicit (learned) geometric model. Jointly, these two models generate grasps that are mechanically fit, and that grip on the object in a way that enables the intended task. Renaud Detry, Jeremie Papon, Larry H. Matthies |
IROS | 2 |
| 2015 | Constrained planar cuts - Object partitioning for point cloudsabstractWhile humans can easily separate unknown objects into meaningful parts, recent segmentation methods can only achieve similar partitionings by training on human-annotated ground-truth data. Here we introduce a bottom-up method for segmenting 3D point clouds into functional parts which does not require supervision and achieves equally good results. Our method uses local concavities as an indicator for inter-part boundaries. We show that this criterion is efficient to compute and generalizes well across different object classes. The algorithm employs a novel locally constrained geometrical boundary model which proposes greedy cuts through a local concavity graph. Only planar cuts are considered and evaluated using a cost function, which rewards cuts orthogonal to concave edges. Additionally, a local clustering constraint is applied to ensure the partitioning only affects relevant locally concave regions. We evaluate our algorithm on recordings from an RGB-D camera as well as the Princeton Segmentation Benchmark, using a fixed set of parameters across all object classes. This stands in stark contrast to most reported results which require either knowing the number of parts or annotated ground-truth for learning. Our approach outperforms all existing bottom-up methods (reducing the gap to human performance by up to 50 %) and achieves scores similar to top-down data-driven approaches. Markus Schoeler, Jeremie Papon, Florentin Wörgötter |
CVPR | 2 |
| 2015 | Semantic Pose Using Deep Networks Trained on Synthetic RGB-DabstractIn this work we address the problem of indoor scene understanding from RGB-D images. Specifically, we propose to find instances of common furniture classes, their spatial extent, and their pose with respect to generalized class models. To accomplish this, we use a deep, wide, multi-output convolutional neural network (CNN) that predicts class, pose, and location of possible objects simultaneously. To overcome the lack of large annotated RGB-D training sets (especially those with pose), we use an on-the-fly rendering pipeline that generates realistic cluttered room scenes in parallel to training. We then perform transfer learning on the relatively small amount of publicly available annotated RGB-D data, and find that our model is able to successfully annotate even highly challenging real scenes. Importantly, our trained network is able to understand noisy and sparse observations of highly cluttered scenes with a remarkable degree of accuracy, inferring class and pose from a very limited set of cues. Additionally, our neural network is only moderately deep and computes class, pose and position in tandem, so the overall run-time is significantly faster than existing methods, estimating all output parameters simultaneously in parallel. Jeremie Papon, Markus Schoeler |
ICCV | 1 |
| 2015 | Spatially Stratified Correspondence Sampling for Real-Time Point Cloud TrackingabstractIn this paper we propose a novel spatially stratified sampling technique for evaluating the likelihood function in particle filters. In particular, we show that in the case where the measurement function uses spatial correspondence, we can greatly reduce computational cost by exploiting spatial structure to avoid redundant computations. We present results which quantitatively show that the technique permits equivalent, and in some cases, greater accuracy, as a reference point cloud particle filter at significantly faster run-times. We also compare to a GPU implementation, and show that we can exceed their performance on the CPU. In addition, we present results on a multi-target tracking application, demonstrating that the increases in efficiency permit online 6DoF multi-target tracking on standard hardware. Jeremie Papon, Markus Schoeler, Florentin Wörgötter |
WACV | 1 |
| 2015 | Unsupervised Generation of Context-Relevant Training-Sets for Visual Object Recognition Employing MultilingualityabstractImage based object classification requires clean training data sets. Gathering such sets is usually done manually by humans, which is time-consuming and laborious. On the other hand, directly using images from search engines creates very noisy data due to ambiguous noun-focused indexing. However, in daily speech nouns and verbs are always coupled. We use this for the automatic generation of clean data sets by the here-presented TRANSCLEAN algorithm, which through the use of multiple languages also solves the problem of polyesters (a single spelling with multiple meanings). Thus, we use the implicit knowledge contained in verbs, e.g. in an imperative such as "hit the nail", implicating a metal nail and not the fingernail. One type of reference application where this method can automatically operate is human-robot collaboration based on discourse. A second is the generation of clean image data sets, where tedious manual cleaning can be replaced by the much simpler manual generation of a single relevant verb-noun tuple. Here we show the impact of our improved training sets for several widely used and state-of-the-art classifiers including Multipath Hierarchical Matching Pursuit. All tested classifiers show a substantial boost of about +20% in recognition performance. Markus Schoeler, Florentin Wörgötter, Tomas Kulvicius, Jeremie Papon |
WACV | 4 |
| 2014 | Object Partitioning Using Local ConvexityabstractThe problem of how to arrive at an appropriate 3D-segmentation of a scene remains difficult. While current state-of-the-art methods continue to gradually improve in benchmark performance, they also grow more and more complex, for example by incorporating chains of classifiers, which require training on large manually annotated data-sets. As an alternative to this, we present a new, efficient learning- and model-free approach for the segmentation of 3D point clouds into object parts. The algorithm begins by decomposing the scene into an adjacency-graph of surface patches based on a voxel grid. Edges in the graph are then classified as either convex or concave using a novel combination of simple criteria which operate on the local geometry of these patches. This way the graph is divided into locally convex connected subgraphs, which -- with high accuracy -- represent object parts. Additionally, we propose a novel depth dependent voxel grid to deal with the decreasing point-density at far distances in the point clouds. This improves segmentation, allowing the use of fixed parameters for vastly different scenes. The algorithm is straightforward to implement and requires no training data, while nevertheless producing results that are comparable to state-of-the-art methods which incorporate high-level concepts involving classification, learning and model fitting. Simon Christoph Stein, Markus Schoeler, Jeremie Papon, Florentin Wörgötter |
CVPR | 3 |
| 2014 | Convexity based object partitioning for robot applicationsabstractThe idea that connected convex surfaces, separated by concave boundaries, play an important role for the perception of objects and their decomposition into parts has been discussed for a long time. Based on this idea, we present a new bottom-up approach for the segmentation of 3D point clouds into object parts. The algorithm approximates a scene using an adjacency-graph of spatially connected surface patches. Edges in the graph are then classified as either convex or concave using a novel, strictly local criterion. Region growing is employed to identify locally convex connected subgraphs, which represent the object parts. We show quantitatively that our algorithm, although conceptually easy to graph and fast to compute, produces results that are comparable to far more complex state-of-the-art methods which use classification, learning and model fitting. This suggests that convexity/concavity is a powerful feature for object partitioning using 3D data. Furthermore we demonstrate that for many objects a natural decomposition into “handle and body” emerges when employing our method. We exploit this property in a robotic application enabling a robot to automatically grasp objects by their handles. Simon Christoph Stein, Florentin Wörgötter, Markus Schoeler, Jeremie Papon, Tomas Kulvicius |
ICRA | 4 |
| 2013 | Voxel Cloud Connectivity Segmentation - Supervoxels for Point CloudsabstractUnsupervised over-segmentation of an image into regions of perceptually similar pixels, known as super pixels, is a widely used preprocessing step in segmentation algorithms. Super pixel methods reduce the number of regions that must be considered later by more computationally expensive algorithms, with a minimal loss of information. Nevertheless, as some information is inevitably lost, it is vital that super pixels not cross object boundaries, as such errors will propagate through later steps. Existing methods make use of projected color or depth information, but do not consider three dimensional geometric relationships between observed data points which can be used to prevent super pixels from crossing regions of empty space. We propose a novel over-segmentation algorithm which uses voxel relationships to produce over-segmentations which are fully consistent with the spatial geometry of the scene in three dimensional, rather than projective, space. Enforcing the constraint that segmented regions must have spatial connectivity prevents label flow across semantic object boundaries which might otherwise be violated. Additionally, as the algorithm works directly in 3D space, observations from several calibrated RGB+D cameras can be segmented jointly. Experiments on a large data set of human annotated RGB+D images demonstrate a significant reduction in occurrence of clusters crossing object boundaries, while maintaining speeds comparable to state-of-the-art 2D methods. Jeremie Papon, Alexey Abramov, Markus Schoeler, Florentin Wörgötter |
CVPR | 1 |
| 2013 | Toward a library of manipulation actions based on semantic object-action relationsabstractThe goal of this study is to provide an architecture for a generic definition of robot manipulation actions. We emphasize that the representation of actions presented here is “procedural”. Thus, we will define the structural elements of our action representations as execution protocols. To achieve this, manipulations are defined using three levels. The toplevel defines objects, their relations and the actions in an abstract and symbolic way. A mid-level sequencer, with which the action primitives are chained, is used to structure the actual action execution, which is performed via the bottom level. This (lowest) level collects data from sensors and communicates with the control system of the robot. This method enables robot manipulators to execute the same action in different situations i.e. on different objects with different positions and orientations. In addition, two methods of detecting action failure are provided which are necessary to handle faults in system. To demonstrate the effectiveness of the proposed framework, several different actions are performed on our robotic setup and results are shown. This way we are creating a library of human-like robot actions, which can be used by higher-level task planners to execute more complex tasks. Mohamad Javad Aein, Eren Erdal Aksoy, Minija Tamosiunaite, Jeremie Papon, Ales Ude, Florentin Wörgötter |
IROS | 4 |
| 2013 | Point cloud video object segmentation using a persistent supervoxel world-modelabstractRobust visual tracking is an essential precursor to understanding and replicating human actions in robotic systems. In order to accurately evaluate the semantic meaning of a sequence of video frames, or to replicate an action contained therein, one must be able to coherently track and segment all observed agents and objects. This work proposes a novel online point cloud based algorithm which simultaneously tracks 6DoF pose and determines spatial extent of all entities in indoor scenarios. This is accomplished using a persistent supervoxel world-model which is updated, rather than replaced, as new frames of data arrive. Maintenance of a world model enables general object permanence, permitting successful tracking through full occlusions. Object models are tracked using a bank of independent adaptive particle filters which use a supervoxel observation model to give rough estimates of object state. These are united using a novel multi-model RANSAC-like approach, which seeks to minimize a global energy function associating world-model supervoxels to predicted states. We present results on a standard robotic assembly benchmark for two application scenarios - human trajectory imitation and semantic action understanding - demonstrating the usefulness of the tracking in intelligent robotic systems. Jeremie Papon, Tomas Kulvicius, Eren Erdal Aksoy, Florentin Wörgötter |
IROS | 1 |
| 2012 | Depth-supported real-time video segmentation with the KinectabstractThis research has received funding by the EU GARNICS project FP7-247947 and the EU IntellAct project FP7-269959. B.Dellen acknowledges support from the Spanish Ministry for Science and Innovation via a Ramon y Cajal fellowship. K. Pauwels acknowledges support from CEI BioTIC \nGENIL (CEB09-0010) of the MICINN CEI program. Alexey Abramov, Karl Pauwels, Jeremie Papon, Florentin Wörgötter, Babette Dellen |
WACV | 3 |
| 2012 | A modular system architecture for online parallel vision pipelinesabstractWe present an architecture for real-time, online vision systems which enables development and use of complex vision pipelines integrating any number of algorithms. Individual algorithms are implemented using modular plugins, allowing integration of independently developed algorithms and rapid testing of new vision pipeline configurations. The architecture exploits the parallelization of graphics processing units (GPUs) and multi-core systems to speed processing and achieve real-time performance. Additionally, the use of a global memory management system for frame buffering permits complex algorithmic flow (e.g. feedback loops) in online processing setups, while maintaining the benefits of threaded asynchronous operation of separate algorithms. To demonstrate the system, a typical real-time system setup is described which incorporates plugins for video and depth acquisition, GPU-based segmentation and optical flow, semantic graph generation, and online visualization of output. Performance numbers are shown which demonstrate the insignificant overhead cost of the architecture as well as speed-up over strictly CPU and single threaded implementations. Jeremie Papon, Alexey Abramov, Eren Erdal Aksoy, Florentin Wörgötter |
WACV | 1 |
| 2012 | Real-Time Segmentation of Stereo Videos on a Portable System With a Mobile GPUabstractIn mobile robotic applications, visual information needs to be processed fast despite resource limitations of the mobile system. Here, a novel real-time framework for model-free spatiotemporal segmentation of stereo videos is presented. It combines real-time optical flow and stereo with image segmentation and runs on a portable system with an integrated mobile graphics processing unit. The system performs online, automatic, and dense segmentation of stereo videos and serves as a visual front end for preprocessing in mobile robots, providing a condensed representation of the scene that can potentially be utilized in various applications, e.g., object manipulation, manipulation recognition, visual servoing. The method was tested on real-world sequences with arbitrary motions, including videos acquired with a moving camera. Alexey Abramov, Karl Pauwels, Jeremie Papon, Florentin Wörgötter, Babette Dellen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |