Özgür Erkent

dblp:83/3473 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
5since 2021 · last 2023
0000-0002-2436-3186ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 3 since 2021Systems, architecture and hardware · 8 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Autonomous driving · 31% Robot navigation and mapping · 29% 3D vision · 20%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Autonomous driving
perception
0.712023
LAPTNet-FPN: Multi-Scale LiDAR-Aided Projective Transform Network for Real Time Semantic Grid Prediction · ICRA 2023
Robotics › Robot navigation and mapping
sensor fusion
0.712023
LAPTNet-FPN: Multi-Scale LiDAR-Aided Projective Transform Network for Real Time Semantic Grid Prediction · ICRA 2023
Robotics › Robot navigation and mapping › robot mapping
topological map
0.422015
Long-term topological place learning · ICRA 2015
Place representation in topological maps based on bubble space · ICRA 2012
Robotics › Autonomous driving
driver behavior modeling
0.312018
Modeling Driver Behavior from Demonstrations in Dynamic Environments Using Spatiotemporal Lattices · ICRA 2018
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning
0.312018
Modeling Driver Behavior from Demonstrations in Dynamic Environments Using Spatiotemporal Lattices · ICRA 2018
Robotics › Motion planning and robot control
trajectory optimization
0.312018
Modeling Driver Behavior from Demonstrations in Dynamic Environments Using Spatiotemporal Lattices · ICRA 2018
Computer vision › 3D vision
camera pose estimation
0.212016
Integration of Probabilistic Pose Estimates from Multiple Views · ECCV (7) 2016
Computer vision › 3D vision › pose estimation
multi-view pose estimation
0.212016
Integration of Probabilistic Pose Estimates from Multiple Views · ECCV (7) 2016
Computer vision › 3D vision › 3d scene understanding
depth and scene understanding
0.212023
LAPTNet-FPN: Multi-Scale LiDAR-Aided Projective Transform Network for Real Time Semantic Grid Prediction · ICRA 2023

Methods — techniques the papers use, named apart from their topics

projective transform network · 0.7point cloud guidance · 0.7feature pyramid network · 0.7inverse reinforcement learning · 0.3conformal spatiotemporal state lattice · 0.3probabilistic pose estimation · 0.2hierarchical single link clustering · 0.2SLINK · 0.2support vector machine · 0.1bubble descriptors · 0.1
YearPublicationVenuePosition
2023 LAPTNet-FPN: Multi-Scale LiDAR-Aided Projective Transform Network for Real Time Semantic Grid Prediction
abstract
Semantic grids can be useful representations of the scene around an autonomous system. By having information about the layout of the space around itself, a robot can leverage this type of representation for crucial tasks such as navigation or tracking. By fusing information from multiple sensors, robustness can be increased and the computational load for the task can be lowered, achieving real time performance. Our multi-scale LiDAR-Aided Perspective Transform network uses information available in point clouds to guide the projection of image features to a top-view representation, resulting in a relative improvement in the state of the art for semantic grid generation for human (+8.67%) and movable object (+49.07%) classes in the nuScenes dataset, as well as achieving results close to the state of the art for the vehicle, drivable area and walkway classes, while performing inference at 25 FPS.
Manuel Diaz-Zapata, David Sierra González, Özgür Erkent, Christian Laugier, Jilles Steeve Dibangoye
ICRA3
2022 LAPTNet: LiDAR-Aided Perspective Transform Network
abstract
Semantic grids are a useful representation of the environment around a robot. They can be used in autonomous vehicles to concisely represent the scene around the car, capturing vital information for downstream tasks like navigation or collision assessment. Information from different sensors can be used to generate these grids. Some methods rely only on RGB images, whereas others choose to incorporate information from other sensors, such as radar or LiDAR. In this paper, we present an architecture that fuses LiDAR and camera information to generate semantic grids. By using the 3D information from a LiDAR point cloud, the LiDAR-Aided Perspective Transform Network (LAPTNet) is able to associate features in the camera plane to the bird's eye view without having to predict any depth information about the scene. Compared to state-of-the-art camera-only methods, LAPTNet achieves an improvement of up to 8.8 points (or 38.13%) over state-of-art competing approaches for the classes proposed in the NuScenes dataset validation split.
Manuel Diaz-Zapata, Özgür Erkent, Christian Laugier, Jilles Steeve Dibangoye, David Sierra González
ICARCV2
2022 TransFuseGrid: Transformer-based Lidar-RGB fusion for semantic grid prediction
abstract
Semantic grids are a succinct and convenient approach to represent the environment for mobile robotics and autonomous driving applications. While the use of Lidar sensors is now generalized in robotics, most semantic grid prediction approaches in the literature focus only on RGB data. In this paper, we present an approach for semantic grid prediction that uses a transformer architecture to fuse Lidar sensor data with RGB images from multiple cameras. Our proposed method, TransFuseGrid, first transforms both input streams into top-view embeddings, and then fuses these embeddings at multiple scales with Transformers. Finally, a decoder transforms the fused, top-view feature map into a semantic grid of the vehicle's environment. We evaluate the performance of our approach on the nuScenes dataset for the vehicle, drivable area, lane divider and walkway segmentation tasks. The results show that Trans-FuseGrid achieves superior performance than competing RGB-only and Lidar-only methods. Additionally, the Transformer feature fusion leads to a significative improvement over naive RGB-Lidar concatenation. In particular, for the segmentation of vehicles, our model outperforms state-of-the-art RGB-only and Lidar-only methods by 24% and 53%, respectively.
Gustavo Salazar-Gomez, David Sierra González, Manuel Diaz-Zapata, Anshul Paigwar, Wenqian Liu, Özgür Erkent, Christian Laugier
ICARCV6
2021 YOLO-based Panoptic Segmentation Network
abstract
Autonomous vehicles need information about their surroundings to safely navigate them. For this, the task of Panoptic Segmentation is proposed as a method of fully parsing the scene by assigning each pixel a label and instance id. Given the constraints of autonomous driving, this process needs to be done in a fast manner. In this paper, we propose the first panoptic segmentation network based on the YOLOv3 real-time object detection network by adding a semantic and instance segmentation branches. YOLO-panoptic is able to do real-time inference and achieves a performance similar to the state of the art methods in some metrics.
Manuel Diaz-Zapata, Özgür Erkent, Christian Laugier
COMPSAC2
2021 GridTrack: Detection and Tracking of Multiple Objects in Dynamic Occupancy Grids
Özgür Erkent, David Sierra González, Anshul Paigwar, Christian Laugier
ICVS1
2020 Instance Segmentation with Unsupervised Adaptation to Different Domains for Autonomous Vehicles
abstract
Detection of the objects around a vehicle is important for a safe and successful navigation of an autonomous vehicle. Instance segmentation provides a fine and accurate classification of the objects such as cars, trucks, pedestrians, etc. In this study, we propose a fast and accurate approach which can detect and segment the object instances which can be adapted to new conditions without requiring the labels from the new condition. Furthermore, the performance of the instance segmentation does not degrade in detection of the objects in the original condition after it adapts to the new condition. To our knowledge, currently there are not other methods which perform unsupervised domain adaptation for the task of instance segmentation using non-synthetic datasets. We evaluate the adaptation capability of our method on two datasets. Firstly, we test its capacity of adapting to a new domain; secondly, we test its ability to adapt to new weather conditions. The results show that it can adapt to new conditions with an improved accuracy while preserving the accuracy of the original condition.
Manuel Diaz-Zapata, Özgür Erkent, Christian Laugier
ICARCV2
2020 Leveraging Dynamic Occupancy Grids for 3D Object Detection in Point Clouds
abstract
Traditionally, point cloud-based 3D object detectors are trained on annotated, non-sequential samples taken from driving sequences (e.g. the KITTI dataset). However, by doing this, the developed algorithms renounce to exploit any dynamic information from the driving sequences. It is reasonable to think that this information, which is available at test time when deploying the models in the experimental vehicles, could have significant predictive potential for the object detection task. To study the advantages that this kind of information could provide, we construct a dataset of dynamic occupancy grid maps from the raw KITTI dataset and find the correspondence to each of the KITTI 3D object detection dataset samples. By training a Lidar-based state-of-the-art 3D object detector with and without the dynamic information we get insights into the predictive value of the dynamics. Our results show that having access to the environment dynamics improves by 27% the ability of the detection algorithm to predict the orientation of smaller obstacles such as pedestrians. Furthermore, the 3D and bird's eye view bounding box predictions for pedestrians in challenging cases also see a 7% improvement. Qualitatively speaking, the dynamics help with the detection of partially occluded and far-away obstacles. We illustrate this fact with numerous qualitative prediction results.
David Sierra González, Anshul Paigwar, Özgür Erkent, Jilles Steeve Dibangoye, Christian Laugier
ICARCV3
2020 Recognize Moving Objects Around an Autonomous Vehicle Considering a Deep-learning Detector Model and Dynamic Bayesian Occupancy
abstract
Perception systems on autonomous vehicles have the challenge of understanding the traffic scene in different situations. The fusion of redundant information obtained from different sources has been shown considerable progress under different methodologies to achieve this objective. However, new opportunities are available to obtain better fusion results with the advance of deep-learning models and computing hardware. In this paper, we aim to recognize moving objects in traffic scenes through the fusion of semantic information with occupancy-grid estimations. Our approach considers a deep-learning model with inference times between 22 to 55 milliseconds. Moreover, we use a Bayesian occupancy framework with a Highly-parallelized design to obtain the occupancy-grid estimations. We validate our approach using experimental results with real-world data on urban scenery.
Andrés E. Gómez Hernandez, Özgür Erkent, Christian Laugier
ICARCV2
2020 GndNet: Fast Ground Plane Estimation and Point Cloud Segmentation for Autonomous Vehicles
abstract
Ground plane estimation and ground point segmentation is a crucial precursor for many applications in robotics and intelligent vehicles like navigable space detection and occupancy grid generation, 3D object detection, point cloud matching for localization and registration for mapping. In this paper, we present GndNet, a novel end-to-end approach that estimates the ground plane elevation information in a grid-based representation and segments the ground points simultaneously in real-time. GndNet uses PointNet and Pillar Feature Encoding network to extract features and regresses ground height for each cell of the grid. We augment the SemanticKITTI dataset to train our network. We demonstrate qualitative and quantitative evaluation of our results for ground elevation estimation and semantic segmentation of point cloud. GndNet establishes a new state-of-the-art, achieves a run-time of 55Hz for ground plane estimation and ground point segmentation.
Anshul Paigwar, Özgür Erkent, David Sierra González, Christian Laugier
IROS2
2018 Semantic Grid Estimation with Occupancy Grids and Semantic Segmentation Networks
abstract
We propose a method to estimate the semantic grid for an autonomous vehicle. The semantic grid is a 2D bird's eye view map where the grid cells contain semantic characteristics such as road, car, pedestrian, signage, etc. We obtain the semantic grid by fusing the semantic segmentation information and an occupancy grid computed by using a Bayesian filter technique. To compute the semantic information from a monocular RGB image, we integrate segmentation deep neural networks into our model. We use a deep neural network to learn the relation between the semantic information and the occupancy grid which can be trained end-to-end extending our previous work on semantic grids. Furthermore, we investigate the effect of using a conditional random field to refine the results. Finally, we test our method on two datasets and compare different architecture types for semantic segmentation. We perform the experiments on KITTI dataset and Inria-Chroma dataset.
Özgür Erkent, Christian Wolf 0001, Christian Laugier
ICARCV1
2018 Modeling Driver Behavior from Demonstrations in Dynamic Environments Using Spatiotemporal Lattices
abstract
One of the most challenging tasks in the development of path planners for intelligent vehicles is the design of the cost function that models the desired behavior of the vehicle. While this task has been traditionally accomplished by hand-tuning the model parameters, recent approaches propose to learn the model automatically from demonstrated driving data using Inverse Reinforcement Learning (IRL). To determine if the model has correctly captured the demonstrated behavior, most IRL methods require obtaining a policy by solving the forward control problem repetitively. Calculating the full policy is a costly task in continuous or large domains and thus often approximated by finding a single trajectory using traditional path-planning techniques. In this work, we propose to find such a trajectory using a conformal spatiotemporal state lattice, which offers two main advantages. First, by conforming the lattice to the environment, the search is focused only on feasible motions for the robot, saving computational power. And second, by considering time as part of the state, the trajectory is optimized with respect to the motion of the dynamic obstacles in the scene. As a consequence, the resulting trajectory can be used for the model assessment. We show how the proposed IRL framework can successfully handle highly dynamic environments by modeling the highway tactical driving task from demonstrated driving data gathered with an instrumented vehicle.
David Sierra González, Özgür Erkent, Victor Romero-Cano, Jilles Steeve Dibangoye, Christian Laugier
ICRA2
2018 Semantic Grid Estimation with a Hybrid Bayesian and Deep Neural Network Approach
abstract
In an autonomous vehicle setting, we propose a method for the estimation of a semantic grid, i.e. a bird's eye grid centered on the car's position and aligned with its driving direction, which contains high-level semantic information about the environment and its actors. Each grid cell contains a semantic label with divers classes, as for instance {Road, Vegetation, Building, Pedestrian, Car...}. We propose a hybrid approach, which combines the advantages of two different methodologies: we use Deep Learning to perform semantic segmentation on monocular RGB images with supervised learning from labeled groundtruth data. We combine these segmentations with occupancy grids calculated from LIDAR data using a generative Bayesian particle filter. The fusion itself is carried out with a deep neural network, which learns to integrate geometric information from the LIDAR with semantic information from the RGB data. We tested our method on two datasets, namely the KITTI dataset, which is publicly available and widely used, and our own dataset obtained with our own platform, equipped with a LIDAR and various sensors. We largely outperform baselines which calculate the semantic grid either from the RGB image alone or from LIDAR output alone, showing the interest of this hybrid approach.
Özgür Erkent, Christian Wolf 0001, Christian Laugier, David Sierra González, Victor Romero-Cano
IROS1
2017 Supervised Learning of Gesture-Action Associations for Human-Robot Collaboration
abstract
As human-robot collaboration methodologies develop robots need to adapt fast learning methods in domestic scenarios. The paper presents a novel approach to learn associations between the human hand gestures and the robot's manipulation actions. The role of the robot is to operate as an assistant to the user. In this context we propose a supervised learning framework to explore the gesture-action space for human-robot collaboration scenario. The framework enables the robot to learn the gesture-action associations on the fly while performing the task with the user; an example of zero-shot learning. We discuss the effect of an accurate gesture detection in performing the task. The accuracy of the gesture detection system directly accounts for the amount of effort put by the user and the number of actions performed by the robot.
Dadhichi Shukla, Özgür Erkent, Justus H. Piater
FG2
2017 Visual task outcome verification using deep learning
abstract
Manipulation tasks requiring high precision are difficult for reasons such as imprecise calibration and perceptual inaccuracies. We present a method for visual task outcome verification that provides an assessment of the task status as well as information for the robot to improve this status. The final status of the task is assessed as success, failure or in progress. We propose a deep learning strategy to learn the task with a small number of training episodes and without requiring the robot. A probabilistic, appearance-based pose estimation method is used to learn the demonstrated task. For real-data efficiency, synthetic training images are created around the trajectory of the demonstrated task. We show that our method can estimate the task status with high accuracy in several instances of different tasks, and demonstrate the accuracy of a high-precision task on a real robot.
Özgür Erkent, Dadhichi Shukla, Justus H. Piater
IROS1
2017 Proactive, incremental learning of gesture-action associations for human-robot collaboration
abstract
Identifying an object of interest, grasping it, and handing it over are key capabilities of collaborative robots. In this context we propose a fast, supervised learning framework for learning associations between human hand gestures and the intended robotic manipulation actions. This framework enables the robot to learn associations on the fly while performing a task with the user. We consider a domestic scenario of assembling a kid's table where the role of the robot is to assist the user. To facilitate the collaboration we incorporate the robot's gaze into the framework. The proposed approach is evaluated in simulation as well as in a real environment. We study the effect of accurate gesture detection on the number of interactions required to complete the task. Moreover, our quantitative analysis shows how purposeful gaze can significantly reduce the amount of time required to achieve the goal.
Dadhichi Shukla, Özgür Erkent, Justus H. Piater
RO-MAN2
2016 Integration of Probabilistic Pose Estimates from Multiple Views
Özgür Erkent, Dadhichi Shukla, Justus H. Piater
ECCV (7)1
2016 A multi-view hand gesture RGB-D dataset for human-robot interaction scenarios
abstract
Understanding semantic meaning from hand gestures is a challenging but essential task in human-robot interaction scenarios. In this paper we present a baseline evaluation of the Innsbruck Multi-View Hand Gesture (IMHG) dataset [1] recorded with two RGB-D cameras (Kinect). As a baseline, we adopt a probabilistic appearance-based framework [2] to detect a hand gesture and estimate its pose using two cameras. The dataset consists of two types of deictic gestures with the ground truth location of the target, two symbolic gestures, two manipulative gestures, and two interactional gestures. We discuss the effect of parallax due to the offset between head and hand while performing deictic gestures. Furthermore, we evaluate the proposed framework to estimate the potential referents on the Innsbruck Pointing at Objects (IPO) dataset [2].
Dadhichi Shukla, Özgür Erkent, Justus H. Piater
RO-MAN2
2015 Long-term topological place learning
abstract
In this work, we consider long-term topological place learning and present an approach that enables the robot to learn in an unsupervised, organized and incremental manner. The knowledge associated with the previously visited places is internally stored in the form of bubble descriptor semantic tree (BDST) using the previously proposed bubble space representation. The BDST is generated and maintained without any external supervision. It organizes the learned knowledge where the terminal nodes are viewed as corresponding to distinct places while its structure encodes their semantic hierarchy. In case the robot is not able to recognize a place with its current BDST, it learns it via updating the BDST incrementally based on the hierarchical single link clustering algorithm SLINK. The proposed approach is evaluated experimentally using combined benchmark datasets from indoor and outdoor settings with recognition rates comparable to those of state-of-the-art approaches while the robot is able to retain efficiently and use the knowledge associated with the learned places.
Özgür Erkent, H. Isil Bozma
ICRA1
2015 General Object Tip Detection and Pose Estimation for Robot Manipulation
Dadhichi Shukla, Özgür Erkent, Justus H. Piater
ICVS2
2014 RGB-D based place representation in topological maps
Hakan Karaoguz, Özgür Erkent, H. Isil Bozma
Mach. Vis. Appl.2
2013 Integrating Cue Descriptors in Bubble Space for Place Recognition
Özgür Erkent, H. Isil Bozma
ICVS1
2012 Place representation in topological maps based on bubble space
abstract
Place representation is a key element in topological maps. This paper presents bubble space - a novel representation for “places” (nodes) in topological maps. The novelties of this model are two-fold: First, a mathematical formalism that defines bubble space is presented. This formalism extends previously proposed bubble memory to accommodate two new variables - varying robot pose and multiple features. Each bubble surface preserves the local S2-metric relations of the incoming sensory data from the robot's viewpoint. Secondly, for learning and recognition, bubble surfaces can be transformed into bubble descriptors that are compact and rotationally invariant, while being computable in an incremental manner. The proposed model is evaluated with support vector machine based decision making in two different settings: first with a mobile robot placed in a variety of locations and secondly using benchmark visual data.
Özgür Erkent, H. Isil Bozma
ICRA1
2006 Saccades and Fixating Using Artificial Potential Functions
abstract
This paper presents a mathematical model for saccadic motion and fixations. We relate this issue to the problem of motion planning and show that a family of artificial potential functions can be used for creating saccadic motion. The advantage of this approach is that finding the next fixation point does not require an explicit visual search - which is computationally costly and may be problematic in real-time applications. Rather, the system naturally 'slides' from the current fixation into the next. Thus real-time performance on cheap hardware can easily be achieved. Experimental results serve to provide insight into the performance of a robot APES implementing this approach
B. Deniz Ilhan, Özgür Erkent, H. Isil Bozma
IROS2