EDBT 2026 Demo / reviewers in the wild / expert
Jana Kosecka
dblp:j/JanaKosecka · also Jana Kosecká
· DBLP profile ↗
64ranked-venue papers
12as first author
7since 2021 · last 2025
0000-0003-4619-3277ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 11 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 2 since 2021Systems, architecture and hardware · 22 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Hybrid Approach to Indoor Social Navigation: Integrating Reactive Local Planning and Proactive Global PlanningabstractWe consider the problem of indoor building-scale social navigation, where the robot must reach a point goal as quickly as possible without colliding with humans who are freely moving around. Factors such as varying crowd densities, unpredictable human behavior, and the constraints of indoor spaces add significant complexity to the navigation task, necessitating a more advanced approach. We propose a modular navigation framework that leverages the strengths of both classical methods and deep reinforcement learning (DRL). Our approach employs a global planner to generate waypoints, assigning soft costs around anticipated pedestrian locations, encouraging caution around potential future positions of humans. Simultaneously, the local planner, powered by DRL, follows these waypoints while avoiding collisions. The combination of these planners enables the agent to perform complex maneuvers and effectively navigate crowded and constrained environments while improving reliability. Many existing studies on social navigation are conducted in simplistic or open environments, limiting the ability of trained models to perform well in complex, real-world settings. To advance research in this area, we introduce a new 2D benchmark designed to facilitate development and testing of social navigation strategies in indoor environments.22Simulator and code: https://github.com/arnabGMU/hybrid_social_nav We benchmark our method against traditional and RL-based navigation strategies, demonstrating that our approach outperforms both. Arnab Debnath, Gregory J. Stein, Jana Kosecka |
ICRA | 3 |
| 2025 | Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
Ivana Benová, Jana Kosecka, Michal Gregor, Martin Tamajka, Marcel Veselý, Marián Simko |
SOFSEM (1) | 2 |
| 2023 | Learning-Augmented Model-Based Planning for Visual ExplorationabstractWe consider the problem of time-limited robotic exploration in previously unseen environments where exploration is limited by a predefined amount of time. We propose a novel exploration approach using learning-augmented model-based planning. We generate a set of sub goals associated with frontiers on the current map and derive a Bellman Equation for exploration with these subgoals. Visual sensing and advances in semantic mapping of indoor scenes are exploited for training a deep convolutional neural network to estimate properties associated with each frontier: the expected unobserved area beyond the frontier and the expected time steps (discretized actions) required to explore it. The proposed model-based planner is guaranteed to explore the whole scene if time permits. We thoroughly evaluate our approach on a large-scale pseudo-realistic indoor dataset (Matterport3D) with the Habitat simulator. We compare our approach with classical and more recent RL-based exploration methods. Our approach surpasses the greedy strategies by 2.1% and the RL-based exploration methods by 8.4% in terms of coverage. Arnab Debnath, Gregory J. Stein, Jana Kosecka |
IROS | 4 |
| 2022 | Message from the Program Chairs: 3DV 2022abstractWe welcome you to the 2022 edition of the International Conference on 3D Vision (3DV 2022). The conference took place in hybrid format: after the hiatus due to the global pandemic, we are happy to be able to host an in-person meeting in Prague, Czech Republic, while attendees who could not travel to Prague still also had the possibility to participate virtually. Angela Dai, Jana Kosecka, Gin Hee Lee, Konrad Schindler |
3DV | 2 |
| 2022 | Object Pose Estimation using Mid-level Visual RepresentationsabstractThis work proposes a novel pose estimation model for object categories that can be effectively transferred to pre-viously unseen environments. The deep convolutional network models (CNN) for pose estimation are typically trained and evaluated on datasets specifically curated for object detection, pose estimation, or 3D reconstruction, which requires large amounts of training data. In this work, we propose a model for pose estimation that can be trained with small amount of data and is built on the top of generic mid-level represen-tations [33] (e.g. surface normal estimation and re-shading). These representations are trained on a large dataset without requiring pose and object annotations. Later on, the predictions are refined with a small CNN neural network that exploits object masks and silhouette retrieval. The presented approach achieves superior performance on the Pix3D dataset [26] and shows nearly 35 % improvement over the existing models when only 25 % of the training data is available. We show that the approach is favorable when it comes to generalization and transfer to novel environments. Towards this end, we introduce a new pose estimation benchmark for commonly encountered furniture categories on challenging Active Vision Dataset [1] and evaluated the models trained on the Pix3D dataset. Negar Nejatishahidin, Pooya Fayyazsanavi, Jana Kosecka |
IROS | 3 |
| 2021 | Improving Sign Video Modeling Using Graph Neural NetworkabstractIn this work, we present an ensemble based sign video recognition method. Our proposed method uses different input representations – such as RGB video and body key-points or pose data – to model sign videos in a multi-modal manner. We represent an input sign video in two ways: the dense frame and the sparse frame inputs. The dense input uses 3D Convolutional Neural Network (CNN) on a 64 frame input window and Long Short Term Memory (LSTM) Network on 32 frame pose input. The sparse input picks 5 representative frames from a sign video, and utilizes CNN and Graph Convolutional Network (GCN) based modeling. These representative frames for a video are selected using pose confidences that are obtained from an off-the-shelf pose estimation model. Our experimental results show that, while the dense 3D CNN model achieves best performance as a single classifier, the GCN based sparse model provides extra recognition capacity. More specifically, the sparse modeling source, when added with the dense modeling in an ensemble manner, can disambiguate similar looking sign classes. Our proposed multi-source ensemble method outperforms several state-of-the-art methods on AUTSL Turkish sign language benchmark dataset. Al Amin Hosain, Huzefa Rangwala, Jana Kosecka |
IEEE BigData | 3 |
| 2021 | Hand Pose Guided 3D Pooling for Word-level Sign Language RecognitionabstractGestures in American Sign Language (ASL) are characterized by fast, highly articulate motion of upper body, including arm movements with complex hand shapes and facial expressions. In this work, we propose a new method for word-level sign recognition from American Sign Language (ASL) using video. Our method uses both motion and hand shape cues while being robust to variations of execution. We exploit the knowledge of the body pose, estimated from an off-the-shelf pose estimator. Using the pose as a guide, we pool spatio-temporal feature maps from different layers of a 3D convolutional neural network. We train separate classifiers using pose guided pooled features from different resolutions and fuse their prediction scores during test time. This leads to a significant improvement in performance on the WLASL benchmark dataset [25]. The proposed approach achieves 10%, 12%, 9.5% and 6.5% performance gain on WLASL100, WLASL300, WLASL1000, WLASL2000 subsets respectively. To demonstrate the robustness of the pose guided pooling and proposed fusion mechanism, we also evaluate our method by fine tuning the model on another dataset. This yields 10% performance improvement for the proposed method using only 0.4% training data during fine tuning stage. Al Amin Hosain, Panneer Selvam Santhalingam, Parth H. Pathak, Huzefa Rangwala, Jana Kosecka |
WACV | 5 |
| 2020 | American Sign Language Recognition Using an FMCW Wireless Sensor (Student Abstract)abstractIn today's digital world, rapid technological advancements continue to lessen the burden of tasks for individuals. Among these tasks is communication across perceived language barriers. Indeed, increased attention has been drawn to American Sign Language (ASL) recognition in recent years. Camera-based and motion detection-based methods have been researched extensively; however, there remains a divide in communication between ASL users and non-users. Therefore, this research team proposes the use of a novel wireless sensor (Frequency-Modulated Continuous-Wave Radar) to help bridge the gap in communication. In short, this device sends out signals that detect the user's body positioning in space. These signals then reflect off the body and back to the sensor, developing thousands of cloud points per second, indicating where the body is positioned in space. These cloud points can then be examined for movement over multiple consecutive time frames using a cell division algorithm, ultimately showing how the body moves through space as it completes a single gesture or sentence. At the end of the project, 95% accuracy was achieved in one-object prediction as well as 80% accuracy on cross-object prediction with 30% other objects' data introduced on 19 commonly used gestures. There are 30 samples for each gesture per person from three persons. Yuanqi Du, Nguyen Dang 0002, Riley Wilkerson, Parth H. Pathak, Huzefa Rangwala, Jana Kosecka |
AAAI | 6 |
| 2020 | Body Pose and Deep Hand-shape Feature Based American Sign Language RecognitionabstractThis work presents an approach for American Sign Language (ASL) gesture recognition from videos. Gestures are comprised of various upper body motions involving hand shapes, motion of both hands with facial expression and head movements. Previous approaches tackled this problem by directly learning 3D convolutional spatio-temporal models from video in a simplified settings with uniform backgrounds. To handle more complex variation in appearance and backgrounds we propose to exploit recent advances in estimation of 2D body pose using Deep Convolutional Neural Networks trained on large corpus of human pose annotations. We use the trajectories of 2D skeletal data estimated from video to train a baseline recursive neural network gesture recognition model. The basic model is further extended using embeddings of hand images obtained from another hand shape recognition model [15] with dynamics modeled by another recursive neural network. The final model learns how to fuse two Long Short Term Model (LSTM) recursive neural network models for skeletal and hand image data. We train and evaluate this model on the GMU-ASL51 dataset of 12 users and 51 ASL gestures [8] demonstrating its superior performance compared to several baseline models. Al Amin Hosain, Panneer Selvam Santhalingam, Parth H. Pathak, Jana Kosecka, Huzefa Rangwala |
DSAA | 4 |
| 2020 | Hierarchical Kinematic Human Mesh Recovery
Georgios Georgakis, Srikrishna Karanam, Terrence Chen, Jana Kosecka, Ziyan Wu 0001 |
ECCV (17) | 5 |
| 2020 | FineHand: Learning Hand Shapes for American Sign Language RecognitionabstractAmerican Sign Language recognition is a difficult gesture recognition problem, characterized by fast, highly articulate gestures. These are comprised of arm movements with different hand shapes, facial expression and head movements. Among these components, hand shape is the vital, often the most discriminative part of a gesture. In this work, we present an approach for effective learning of hand shape embeddings, which are discriminative for ASL gestures. For hand shape recognition our method uses a mix of manually labelled hand shapes and high confidence predictions to train deep convolutional neural network (CNN). The sequential gesture component is captured by recursive neural network (RNN) trained on the embeddings learned in the first stage. We will demonstrate that higher quality hand shape models can significantly improve the accuracy of final video gesture classification in challenging conditions with variety of speakers, different illumination and significant motion blurr. We compare our model to alternative approaches exploiting different modalities and representations of the data and show improved video gesture recognition accuracy on GMU-ASL51 benchmark dataset. Al Amin Hosain, Panneer Selvam Santhalingam, Parth H. Pathak, Huzefa Rangwala, Jana Kosecka |
FG | 5 |
| 2020 | Learning View and Target Invariant Visual Servoing for NavigationabstractThe advances in deep reinforcement learning recently revived interest in data-driven learning based approaches to navigation. In this paper we propose to learn viewpoint invariant and target invariant visual servoing for local mobile robot navigation; given an initial view and the goal view or an image of a target, we train deep convolutional network controller to reach the desired goal. We present a new architecture for this task which rests on the ability of establishing correspondences between the initial and goal view and novel reward structure motivated by the traditional feedback control error. The advantage of the proposed model is that it does not require calibration and depth information and achieves robust visual servoing in a variety of environments and targets without any parameter fine tuning. We present comprehensive evaluation of the approach and comparison with other deep learning architectures as well as classical visual servoing methods in visually realistic simulation environment [1]. The presented model overcomes the brittleness of classical visual servoing based methods and achieves significantly higher generalization capability compared to the previous learning approaches. Jana Kosecka |
ICRA | 2 |
| 2019 | Sign Language Recognition Analysis using Multimodal DataabstractVoice-controlled personal and home assistants (such as the Amazon Echo and Apple Siri) are becoming increasingly popular for a variety of applications. However, the benefits of these technologies are not readily accessible to Deaf or Hard-of-Hearing (DHH) users. The objective of this study is to develop and evaluate a sign recognition system using multiple modalities that can be used by DHH signers to interact with voice-controlled devices. With the advancement of depth sensors, skeletal data is used for applications like video analysis and activity recognition. Despite having similarity with the well-studied human activity recognition, the use of 3D skeleton data in sign language recognition is rare. This is because unlike activity recognition, sign language is mostly dependent on hand shape pattern. In this work, we investigate the feasibility of using skeletal and RGB video data for sign language recognition using a combination of different deep learning architectures. We validate our results on a large-scale American Sign Language (ASL) dataset of 12 users and 13107 samples across 51 signs. It is named as GMU-ASL51. We collected the dataset over 6 months and it will be publicly released in the hope of spurring further machine learning research towards providing improved accessibility for digital assistants. Al Amin Hosain, Panneer Selvam Santhalingam, Parth H. Pathak, Jana Kosecka, Huzefa Rangwala |
DSAA | 4 |
| 2019 | Learning Local RGB-to-CAD Correspondences for Object Pose EstimationabstractWe consider the problem of 3D object pose estimation. While much recent work has focused on the RGB domain, the reliance on accurately annotated images limits generalizability and scalability. On the other hand, the easily available object CAD models are rich sources of data, providing a large number of synthetically rendered images. In this paper, we solve this key problem of existing methods requiring expensive 3D pose annotations by proposing a new method that matches RGB images to CAD models for object pose estimation. Our key innovations compared to existing work include removing the need for either real-world textures for CAD models or explicit 3D pose annotations for RGB images. We achieve this through a series of objectives that learn how to select keypoints and enforce viewpoint and modality invariance across RGB images and CAD model renderings. Our experiments demonstrate that the proposed method can reliably estimate object pose in RGB images and generalize to object instances not seen during training. Georgios Georgakis, Srikrishna Karanam, Ziyan Wu 0001, Jana Kosecka |
ICCV | 4 |
| 2019 | Visual Representations for Semantic Target Driven NavigationabstractWhat is a good visual representation for navigation? We study this question in the context of semantic visual navigation, which is the problem of a robot finding its way through a previously unseen environment to a target object, e.g. go to the refrigerator. Instead of acquiring a metric semantic map of an environment and using planning for navigation, our approach learns navigation policies on top of representations that capture spatial layout and semantic contextual cues. We propose to use semantic segmentation and detection masks as observations obtained by state-of-the-art computer vision algorithms and use a deep network to learn the navigation policy. The availability of equitable representations in simulated environments enables joint training using real and simulated data and alleviates the need for domain adaptation or domain randomization commonly used to tackle the sim-to-real transfer of the learned policies. Both the representation and the navigation policy can be readily applied to real non-synthetic environments as demonstrated on the Active Vision Dataset [1]. Our approach successfully gets to the target in 54% of the cases in unexplored environments, compared to 46% for a non-learning based approach, and 28% for a learning-based baseline. Arsalan Mousavian, Alexander Toshev, Marek Fiser, Jana Kosecka, Ayzaan Wahid, James Davidson |
ICRA | 4 |
| 2018 | End-to-End Learning of Keypoint Detector and Descriptor for Pose Invariant 3D MatchingabstractFinding correspondences between images or 3D scans is at the heart of many computer vision and image retrieval applications and is often enabled by matching local keypoint descriptors. Various learning approaches have been applied in the past to different stages of the matching pipeline, considering detection, description, or metric learning objectives. These objectives were typically addressed separately and most previous work has focused on image data. This paper proposes an end-to-end learning framework for keypoint detection and its representation (descriptor) for 3D depth maps or 3D scans, where the two can be jointly optimized towards task-specific objectives without a need for separate annotations. We employ a Siamese architecture augmented by a sampling layer and a novel score loss function which in turn affects the selection of region proposals. The positive and negative examples are obtained automatically by sampling corresponding region proposals based on their consistency with known 3D pose labels. Matching experiments with depth data on multiple benchmark datasets demonstrate the efficacy of the proposed approach, showing significant improvements over state-of-the-art methods. Georgios Georgakis, Srikrishna Karanam, Ziyan Wu 0001, Jan Ernst, Jana Kosecka |
CVPR | 5 |
| 2018 | FarSight: Long-Range Depth Estimation from Outdoor ImagesabstractThis paper introduces the problem of long-range monocular depth estimation for outdoor urban environments. Range sensors and traditional depth estimation algorithms (both stereo and single view) predict depth for distances of less than 100 meters in outdoor settings and 10 meters in indoor settings. The shortcomings of outdoor single view methods that use learning approaches are, to some extent, due to the lack of long-range ground truth training data, which in turn is due to limitations of range sensors. To circumvent this, we first propose a novel strategy for generating synthetic long-range ground truth depth data. We utilize Google Earth images to reconstruct large-scale 3D models of different cities with proper scale. The acquired repository of 3D models and associated RGB views along with their long-range depth renderings are used as training data for depth prediction. We then train two deep neural network models for long-range depth estimation: i) a Convolutional Neural Network (CNN) and ii) a Generative Adversarial Network (GAN). We found in our experiments that the GAN model predicts depth more accurately. We plan to open-source the database and the baseline models for public use. Md. Alimoor Reza, Jana Kosecka, Philip David |
IROS | 2 |
| 2017 | 3D Bounding Box Estimation Using Deep Learning and GeometryabstractWe present a method for 3D object detection and pose estimation from a single image. In contrast to current techniques that only regress the 3D orientation of an object, our method first regresses relatively stable 3D object properties using a deep convolutional neural network and then combines these estimates with geometric constraints provided by a 2D object bounding box to produce a complete 3D bounding box. The first network output estimates the 3D object orientation using a novel hybrid discrete-continuous loss, which significantly outperforms the L2 loss. The second output regresses the 3D object dimensions, which have relatively little variance compared to alternatives and can often be predicted for many object types. These estimates, combined with the geometric constraints on translation imposed by the 2D bounding box, enable us to recover a stable and accurate 3D object pose. We evaluate our method on the challenging KITTI object detection benchmark [2] both on the official metric of 3D orientation estimation and also on the accuracy of the obtained 3D bounding boxes. Although conceptually simple, our method outperforms more complex and computationally expensive approaches that leverage semantic segmentation, instance level segmentation and flat ground priors [4] and sub-category detection [23][24]. Our discrete-continuous loss also produces state of the art results for 3D viewpoint estimation on the Pascal 3D+ dataset[26]. Arsalan Mousavian, Dragomir Anguelov, John Flynn, Jana Kosecka |
CVPR | 4 |
| 2017 | A dataset for developing and benchmarking active visionabstractWe present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured in 9 unique scenes. We train a fast object category detector for instance detection on our data. Using the dataset we show that, although increasingly accurate and fast, the state of the art for object detection is still severely impacted by object scale, occlusion, and viewing direction all of which matter for robotics applications. We next validate the dataset for simulating active vision, and use the dataset to develop and evaluate a deep-network-based system for next best move prediction for object classification using reinforcement learning. Our dataset is available for download at cs.unc.edu/~ammirato/active_vision_dataset_website/. Phil Ammirato, Patrick Poirson, Eunbyung Park, Jana Kosecka, Alexander C. Berg |
ICRA | 4 |
| 2017 | Dense piecewise planar RGB-D SLAM for indoor environmentsabstractThe paper exploits weak Manhattan constraints to parse the structure of indoor environments from RGB-D video sequences in an online setting. We extend the previous approach for single view parsing of indoor scenes to video sequences and formulate the problem of recovering the floor plan of the environment as an optimal labeling problem solved using dynamic programming. The temporal continuity is enforced in a recursive setting, where labeling from previous frames is used as a prior term in the objective function. In addition to recovery of piecewise planar weak Manhattan structure of the extended environment, the orthogonality constraints are also exploited by visual odometry and pose graph optimization. This yields reliable estimates in the presence of large motions and absence of distinctive features to track. We evaluate our method on several challenging indoors sequences demonstrating accurate SLAM and dense mapping of low texture environments. On existing TUM benchmark [19] we achieve competitive results with the alternative approaches which fail in our environments. Phi-Hung Le, Jana Kosecka |
IROS | 2 |
| 2017 | Label propagation in RGB-D videoabstractWe propose a new method for the propagation of semantic labels in RGB-D video of indoor scenes given a set of ground truth keyframes. Manual labeling of all pixels in every frame of a video sequence is labor intensive and costly, yet required for training and testing of semantic segmentation methods. The availability of video enables propagation of labels between the frames for obtaining a large amounts of annotated pixels. While previous methods commonly used optical flow motion cues for label propagation, we present a novel approach using the camera poses and 3D point clouds for propagating the labels in superpixels computed on the unannotated frames of the sequence. The propagation task is formulated as an energy minimization problem in a Conditional Random Field (CRF). We performed experiments on 8 video sequences from SUN3D dataset [1] and showed superior performance to an optical flow based label propagation approach. Furthermore, we demonstrated that the propagated labels can be used to learn better models using data hungry deep convolutional neural network (DCNN) based approaches for the task of semantic segmentation. The approach demonstrates an increase in performance when the ground truth keyframes are combined with the propagated labels during training. Md. Alimoor Reza, Georgios Georgakis, Jana Kosecka |
IROS | 4 |
| 2016 | Multiview RGB-D Dataset for Object Instance DetectionabstractThis paper presents a new multi-view RGB-D dataset of nine kitchen scenes, each containing several objects in realistic cluttered environments including a subset of objects from the BigBird dataset. The viewpoints of the scenes are densely sampled and objects in the scenes are annotated with bounding boxes and in the 3D point cloud. Also, an approach for detection and recognition is presented, which is comprised of two parts: (i) a new multi-view 3D proposal generation method and (ii) the development of several recognition baselines using AlexNet to score our proposals, which is trained either on crops of the dataset or on synthetically composited training images. Finally, we compare the performance of the object proposals and a detection baseline to the Washington RGB-D Scenes (WRGB-D) dataset and demonstrate that our Kitchen scenes dataset is more challenging for object detection and recognition. The dataset is available at: http://cs.gmu.edu/~robot/gmu-kitchens.html. Georgios Georgakis, Md. Alimoor Reza, Arsalan Mousavian, Phi-Hung Le, Jana Kosecka |
3DV | 5 |
| 2016 | Joint Semantic Segmentation and Depth Estimation with Deep Convolutional NetworksabstractMulti-scale deep CNNs have been used successfully for problems mapping each pixel to a label, such as depth estimation and semantic segmentation. It has also been shown that such architectures are reusable and can be used for multiple tasks. These networks are typically trained independently for each task by varying the output layer(s) and training objective. In this work we present a new model for simultaneous depth estimation and semantic segmentation from a single RGB image. Our approach demonstrates the feasibility of training parts of the model for each task and then fine tuning the full, combined model on both tasks simultaneously using a single loss function. Furthermore we couple the deep CNN with fully connected CRF, which captures the contextual relationships and interactions between the semantic and depth cues improving the accuracy of the final results. The proposed model is trained and evaluated on NYUDepth V2 dataset [23] outperforming the state of the art methods on semantic segmentation and achieving comparable results on the task of depth estimation. Arsalan Mousavian, Hamed Pirsiavash, Jana Kosecka |
3DV | 3 |
| 2016 | Fast Single Shot Detection and Pose EstimationabstractFor applications in navigation and robotics, estimating the 3D pose of objects is as important as detection. Many approaches to pose estimation rely on detecting or tracking parts or keypoints [11, 21]. In this paper we build on a recent state-of-the-art convolutional network for sliding-window detection [10] to provide detection and rough pose estimation in a single shot, without intermediate stages of detecting parts or initial bounding boxes. While not the first system to treat pose estimation as a categorization problem, this is the first attempt to combine detection and pose estimation at the same level using a deep learning approach. The key to the architecture is a deep convolutional network where scores for the presence of an object category, the offset for its location, and the approximate pose are all estimated on a regular grid of locations in the image. The resulting system is as accurate as recent work on pose estimation (42.4% 8 View mAVP on Pascal 3D+ [21] ) and significantly faster (46 frames per second (FPS) on a TITAN X GPU). This approach to detection and rough pose estimation is fast and accurate enough to be widely applied as a pre-processing step for tasks including high-accuracy pose estimation, object tracking and localization, and vSLAM. Patrick Poirson, Phil Ammirato, Cheng-Yang Fu, Wei Liu 0015, Jana Kosecka, Alexander C. Berg |
3DV | 5 |
| 2016 | RGB-D multi-view object detection with object proposals and shape contextabstractWe propose a novel approach for multi-view object detection in 3D scenes reconstructed from RGB-D sensor. We utilize shape based representation using local shape context descriptors along with the voting strategy which is supported by unsupervised object proposals generated from 3D point cloud data. Our algorithm starts with a single-view object detection where object proposals generated in 3D space are combined with object specific hypotheses generated by the voting strategy. To tackle the multi-view setting, the data association between multiple views enabled view registration and 3D object proposals. The evidence from multiple views is combined in simple bayesian setting. The approach is evaluated on the Washington RGB-D scenes datasets [1], [2] containing several classes of objects in a table top setting. We evaluated our approach against the other state-of-the-art methods and demonstrated superior performance on the same dataset. Georgios Georgakis, Md. Alimoor Reza, Jana Kosecka |
IROS | 3 |
| 2015 | Semantically guided location recognition for outdoors scenesabstractThe problem of image based localization has a long history both in robotics and computer vision and shares many similarities with image based retrieval problem. Existing techniques use either local features or (semi)-global image signatures in the context of topological mapping or loop closure detection. Difficulties of the location recognition problem are often affected by large appearance and viewpoint variation between the query view and reference dataset and presence of non-discriminative features due to vegetation, sky and road. In this work we show that semantic segmentation labeling of man-made structures can inform the traditional bag-of-visual words models to obtain proper feature weighting and improve the overall location recognition accuracy. We also demonstrate additional capability of identifying individual buildings and estimating their extent in images, providing the essential building block for semantic localization. Towards this end we introduce a new challenging outdoors urban dataset exhibiting large variations in appearance and viewpoint. Arsalan Mousavian, Jana Kosecka, Jyh-Ming Lien |
ICRA | 2 |
| 2014 | Semantic segmentation with heterogeneous sensor coveragesabstractWe propose a new approach to semantic parsing, which can seamlessly integrate evidence from multiple sensors with overlapping but possibly different fields of view (FOV), account for missing data and predict semantic labels over the spatial union of sensors coverages. The existing approaches typically carry out semantic segmentation using only one modality, incorrectly interpolate measurements of other modalities or at best assign semantic labels only to the spatial intersection of coverages of different sensors. In this work we remedy these problems by proposing an effective and efficient strategy for inducing the graph structure of Conditional Random Field used for inference and a novel method for computing the sensor domain dependent potentials. We focus on RGB cameras and 3D data from lasers or depth sensors. The proposed approach achieves superior performance, compared to state of the art and obtains labels for the union of spatial coverages of both sensors, while effectively using appearance or 3D cues when they are available. The efficiency of the approach is amenable to realtime implementation. We quantitatively validate our proposal in two publicly available datasets from indoors and outdoors real environments. The obtained semantic understanding of the acquired sensory information can enable higher level tasks for autonomous mobile robots and facilitate semantic mapping of the environments. Cesar Dario Cadena Lerma, Jana Kosecka |
ICRA | 2 |
| 2014 | Introspective semantic segmentationabstractTraditional approaches for semantic segmentation work in a supervised setting assuming a fixed number of semantic categories and require sufficiently large training sets. The performance of various approaches is often reported in terms of average per pixel class accuracy and global accuracy of the final labeling. When applying the learned models in the practical settings on large amounts of unlabeled data, possibly containing previously unseen categories, it is important to properly quantify their performance by measuring a classifier's introspective capability. We quantify the confidence of the region classifiers in the context of a non-parametric k-nearest neighbor (k-NN) framework for semantic segmentation by using the so called strangeness measure. The proposed measure is evaluated by introducing confidence based image ranking and showing its feasibility on a dataset containing a large number of previously unseen categories. Gautam Singh, Jana Kosecka |
WACV | 2 |
| 2013 | Nonparametric Scene Parsing with Adaptive Feature Relevance and Semantic ContextabstractThis paper presents a nonparametric approach to semantic parsing using small patches and simple gradient, color and location features. We learn the relevance of individual feature channels at test time using a locally adaptive distance metric. To further improve the accuracy of the nonparametric approach, we examine the importance of the retrieval set used to compute the nearest neighbours using a novel semantic descriptor to retrieve better candidates. The approach is validated by experiments on several datasets used for semantic parsing demonstrating the superiority of the method compared to the state of art approaches. Gautam Singh, Jana Kosecka |
CVPR | 2 |
| 2013 | Recursive Inference for Prediction of Objects in Urban Environments
Cesar Dario Cadena Lerma, Jana Kosecka |
ISRR | 2 |
| 2013 | Localization in Urban Environments Using a Panoramic Gist DescriptorabstractVision-based topological localization and mapping for autonomous robotic systems have received increased research interest in recent years. The need to map larger environments requires models at different levels of abstraction and additional abilities to deal with large amounts of data efficiently. Most successful approaches for appearance-based localization and mapping with large datasets typically represent locations using local image features. We study the feasibility of performing these tasks in urban environments using global descriptors instead and taking advantage of the increasingly common panoramic datasets. This paper describes how to represent a panorama using the global gist descriptor, while maintaining desirable invariance properties for location recognition and loop detection. We propose different gist similarity measures and algorithms for appearance-based localization and an online loop-closure detection method, where the probability of loop closure is determined in a Bayesian filtering framework using the proposed image representation. The extensive experimental validation in this paper shows that their performance in urban environments is comparable with local-feature-based approaches when using wide field-of-view images. Ana Cristina Murillo, Gautam Singh, Jana Kosecka, Josechu J. Guerrero |
IEEE Trans. Robotics | 3 |
| 2012 | Detecting Changes in Images of Street Scenes
Jana Kosecka |
ACCV (4) | 1 |
| 2012 | Acquiring semantics induced topology in urban environmentsabstractMethods for acquisition and maintenance of an environment model are central to a broad class of mobility and navigation problems. Towards this end, various metric, topological or hybrid models have been proposed. Due to recent advances in sensing and recognition, acquisition of semantic models of the environments have gained increased interest in the community. In this work, we will demonstrate a capability of using weak semantic models of the environment to induce different topological models, capturing the spatial semantics of the environment at different levels. In the first stage of the model acquisition, we propose to compute semantic layout of the street scenes imagery by recognizing and segmenting buildings, roads, sky, cars and trees. Given such semantic layout, we propose an informative feature characterizing the layout and train a classifier to recognize street intersections in challenging urban inner city scenes. We also show how the evidence of different semantic concepts can induce useful topological representation of the environment, which can aid navigation and localization tasks. To demonstrate the approach, we carry out experiments on a challenging dataset of omnidirectional inner city street views and report the performance of both semantic segmentation and intersection classification. Gautam Singh, Jana Kosecka |
ICRA | 2 |
| 2012 | Special issue on Virtual Representations and Modeling of Large-scale environments (VRML)
Jan-Michael Frahm, Marc Pollefeys, Frank Dellaert, Jana Kosecka |
Comput. Vis. Image Underst. | 4 |
| 2011 | Label propagation in videos indoors with an incremental non-parametric model updateabstractSemantic interpretation of the environment can significantly improve the capabilities of our autonomous robots. This work is focused on automatic semantic label propagation in video of indoor environments acquired by a mobile robot. Using a small number of training examples, we propose a new approach to recognize and label dominant background regions of interest, such as floor, wall and doors, and separate them from the remaining of foreground/object image categories. Our approach performs the labeling at the level of image superpixels. A simple non-parametric model is initialized from a few hand labeled examples in the first frame, and then it is propagated and updated along the sequence. We demonstrate the promising results obtained with our proposal in five different indoor sequences from different environments. The obtained semantic labeling can be used both for autonomous navigation and to provide better context for subsequent object detection. J. Rituerto, Ana Cristina Murillo, Jana Kosecka |
IROS | 3 |
| 2010 | Multi-view Superpixel Stereo in Urban Environments
Branislav Micusík, Jana Kosecka |
Int. J. Comput. Vis. | 2 |
| 2009 | Piecewise planar city 3D modeling from street view panoramic sequencesabstractCity environments often lack textured areas, contain repetitive structures, strong lighting changes and therefore are very difficult for standard 3D modeling pipelines. We present a novel unified framework for creating 3D city models which overcomes these difficulties by exploiting image segmentation cues as well as presence of dominant scene orientations and piecewise planar structures. Given panoramic street view sequences, we first demonstrate how to robustly estimate camera poses without a need for bundle adjustment and propose a multi-view stereo method which operates directly on panoramas, while enforcing the piecewise planarity constraints in the sweeping stage. At last, we propose a new depth fusion method which exploits the constraints of urban environments and combines advantages of volumetric and viewpoint based fusion methods. Our technique avoids expensive voxelization of space, operates directly on 3D reconstructed points through effective kd-tree representation, and obtains a final surface by tessellation of backprojections of those points into the reference image. Branislav Micusík, Jana Kosecka |
CVPR | 2 |
| 2008 | Detection and matching of rectilinear structuresabstractIndoor and outdoor urban environments posses many regularities which can be efficiently exploited and used for general image parsing tasks. We present a novel approach for detecting rectilinear structures and demonstrate their use for wide baseline stereo matching, planar 3D reconstruction, and computation of geometric context. Assuming a presence of dominant orthogonal vanishing directions, we proceed by formulating the detection of the rectilinear structures as a labeling problem on detected line segments. The line segment labels, respecting the proposed grammar rules, are established as the MAP assignment of the corresponding MRF. The proposed framework allows to detect both full as well as partial rectangles, rectangle-in-rectangle structures, and rectangles sharing edges. The use of detected rectangles is demonstrated in the context of difficult wide baseline matching tasks in the presence of repetitive structures and large appearance changes. Branislav Micusík, Horst Wildenauer, Jana Kosecka |
CVPR | 3 |
| 2008 | Motion bias and structure distortion induced by intrinsic calibration errors
Marco Zucchelli, Jana Kosecka |
Image Vis. Comput. | 2 |
| 2007 | Hierarchical building recognition
Wei Zhang 0018, Jana Kosecka |
Image Vis. Comput. | 2 |
| 2007 | Estimating Planar Surface Orientation Using Bispectral AnalysisabstractIn this correspondence, we propose a direct method for estimating the orientation of a plane from a single view under perspective projection. Assuming that the underlying planar texture has random phase, we show that the nonlinearities introduced by perspective projection lead to higher order correlations in the frequency domain. We also empirically show that these correlations are proportional to the orientation of the plane. Minimization of these correlations, using tools from polyspectral analysis, yields the orientation of the plane. We show the efficacy of this technique on synthetic and natural images. Hany Farid, Jana Kosecka |
IEEE Trans. Image Process. | 2 |
| 2006 | Probabilistic Location Recognition using Reduced Feature SetabstractThe localization capability is central to basic navigation tasks and motivates development of various visual navigation systems. In this paper we describe a two stage approach for localization in indoor environments. In the first stage, the environment is partitioned into several locations, each characterized by a set of scale-invariant keypoints and their associated descriptors. In the second stage the keypoints of the query view are integrated probabilistically yielding an estimate of most likely location. The novelty of our approach is in the selection of discriminative features, best suited for characterizing individual locations. We demonstrate that high location recognition rate is maintained with only 10% of the originally detected features, yielding a substantial speedup in recognition and capability of handling larger environments. The ambiguities due to the self-similarity and dynamic changes in the environment are resolved by exploiting spatial relationships between locations captured by hidden Markov model Fayin Li, Jana Kosecka |
ICRA | 2 |
| 2005 | Extraction, matching, and pose recovery based on dominant rectangular structures
Jana Kosecka, Wei Zhang 0018 |
Comput. Vis. Image Underst. | 1 |
| 2004 | Vision based Topological Markov LocalizationabstractIn this paper we study the problem of acquiring a topological model of indoors environment by means of visual sensing and subsequent localization given the model. The resulting model consists of a set of locations and neighborhood relationships between them. Each location in the model is represented by a collection of representative views and their associated descriptors selected from a temporally sub-sampled video stream captured by a mobile robot during exploration. We compare the recognition performance using global image histograms as well as local scale-invariant features as image descriptors, demonstrate their strengths and weaknesses and show how to model the spatial relationships between individual locations by a Hidden Markov Model. The quality of the acquired model is tested in the localization stage by means of location recognition: given a new view or a sequence of views, the most likely location where that view came from is determined. Jana Kosecka, Fayin Li |
ICRA | 1 |
| 2004 | Rank Conditions on the Multiple-View Matrix
Yi Ma 0001, Kun Huang 0001, René Vidal, Jana Kosecka, S. Shankar Sastry |
Int. J. Comput. Vis. | 4 |
| 2003 | Qualitative Image Based Localization in Indoors EnvironmentsabstractMan made indoor environments possess regularities, which can be efficiently exploited in automated model acquisition by means of visual sensing. In this context we propose an approach for inferring a topological model of an environment from images or the video stream captured by a mobile robot during exploration. The proposed model consists of a set of locations and neighborhood relationships between them. Initially each location in the model is represented by a collection of similar, temporally adjacent views, with the similarity defined according to a simple appearance based distance measure. The sparser representation is obtained in a subsequent learning stage by means of learning vector quantization (LVQ). The quality of the model is tested in the context of qualitative localization scheme by means of location recognition: given a new view, the most likely location where that view came from is determined. Jana Kosecka, Philip Barber, Zoran Duric |
CVPR (2) | 1 |
| 2003 | Communication enhanced navigation strategies for teams of mobile agentsabstractIn multi-agent systems engaged in cooperative activities there is an apparent trade-off between the complexity of the individual agents, their sensing capabilities and communication required for accomplishment of particular tasks. One of the main computationally intensive components which affects the complexity of the overall system is the acquisition and maintenance of the environment model where the agents reside. In this paper, in the context of foraging and coordinated traversal task, we will examine control strategies that in the absence of the global model of the environment can substantially improve the performance of the team using additional sensing and communication capabilities. In one case the coordinated strategy is motivated by an ant trail following behavior while in another case the line of sight information is used to constrain the movement of individual agents guaranteeing shorter total traversal times. Justin Hayes, Martha McJunkin, Jana Kosecka |
IROS | 3 |
| 2002 | Video Compass
Jana Kosecka, Wei Zhang 0018 |
ECCV (4) | 1 |
| 2002 | Efficient Computation of Vanishing PointsabstractMan-made environments possess a lot of regularities which simplify otherwise difficult pose estimation and visual reconstruction tasks. The constraints arising front parallel and orthogonal lines and planes can be efficiently exploited at various stages of vision processing pipeline. In this paper we propose an approach for estimation of vanishing points by exploiting the constraints of structured man-made environments, where the majority of lines is aligned with the principal orthogonal directions of the world coordinate frame. We combine efficient image processing techniques used in the line detection and initialization stage with simultaneous grouping and estimation of vanishing directions using expectation maximization (EM) algorithm. Since we assume an uncalibrated camera the estimated vanishing points can be used towards partial camera calibration and estimation of the relative orientation of the camera with respect to the scene. The presented approach is computationally efficient and has been verified extensively by experiments. Jana Kosecka, Wei Zhang 0018 |
ICRA | 1 |
| 2001 | Motion bias and structure distortion induced by calibration errorsabstractThis article provides an account of sensitivity and robustness of structure and motion recovery with respect to the errors in intrinsic parameters of the camera. We demonstrate both analytically and in simulation, the interplay between measurement and calibration errors and their effect on motion and structure estimates. In particular we show that the calibration errors introduce an additional bias towards the optical axis, which has opposite sign to the bias typically observed by egomotion algorithms. The overall bias causes a distortion of the resulting 3D structure, which we express in a parametric form. The analysis and experiments are carried out in the differential setting for motion and structure estimation from image velocities. While the analytical explanations are derived in the context of linear techniques for motion estimation, we verify our observations experimentally on a variety of optimal and suboptimal motion and structure estimation algorithms. The obtained results illuminate and explain the performance and sensitivity of the differential structure and motion recovery techniques in the presence of calibration errors. 1 Marco Zucchelli, Jana Kosecka |
BMVC | 2 |
| 2001 | Optimization Criteria and Geometric Algorithms for Motion and Structure Estimation
Yi Ma 0001, Jana Kosecka, S. Shankar Sastry |
Int. J. Comput. Vis. | 2 |
| 2000 | Kruppa Equation Revisited: Its Renormalization and Degeneracy
Yi Ma 0001, René Vidal, Jana Kosecka, S. Shankar Sastry |
ECCV (2) | 3 |
| 2000 | Hierarchies of Sensing and Control in Visually Guided Agents
Jana Kosecka |
SOFSEM | 1 |
| 2000 | Linear Differential Algorithm for Motion Recovery: A Geometric Approach
Yi Ma 0001, Jana Kosecka, S. Shankar Sastry |
Int. J. Comput. Vis. | 2 |
| 2000 | Euclidean Reconstruction and Reprojection Up to Subgroups
Yi Ma 0001, Stefano Soatto, Jana Kosecka, S. Shankar Sastry |
Int. J. Comput. Vis. | 3 |
| 1999 | Euclidean Reconstruction and Reprojection up to SubgroupsabstractThe necessary and sufficient conditions for being able to estimate scene structure, motion and camera calibration from a sequence of images are very rarely satisfied in practice. What exactly can be estimated in sequences of practical importance, when such conditions are not satisfied? In this paper we give a complete answer to this question. For every camera motion that fails to meet the conditions, we give explicit formulas for the ambiguities in the reconstructed scene, motion and calibration. Such a characterization is crucial both for designing robust estimation algorithms (that do not try to recover parameters that cannot be recovered), and for generating novel views of the scene by controlling the vantage point. To this end, we characterize explicitly all the vantage points that give rise to a valid Euclidean reprojection regardless of the ambiguity in the reconstruction. We also characterize vantage points that generate views that are altogether invariant to the ambiguity. All the results are presented using simple notation that involves no tensors nor complex projective geometry, and should be accessible with basic background in linear algebra. Yi Ma 0001, Stefano Soatto, Jana Kosecka, S. Shankar Sastry |
ICCV | 3 |
| 1999 | Vision guided navigation for a nonholonomic mobile robotabstractTheoretical and analytical aspects of the visual servoing problem have not received much attention. Furthermore, the problem of estimation from the vision measurements has been considered separately from the design of the control strategies. Instead of addressing the pose estimation and control problems separately, we attempt to characterize the types of control tasks which can be achieved using only quantities directly measurable in the image, bypassing the pose estimation phase. We consider the task of navigation for a nonholonomic ground mobile base tracking an arbitrarily shaped continuous ground curve. This tracking problem is formulated as one of controlling the shape of the curve in the image plane. We study the controllability of the system characterizing the dynamics of the image curve, and show that the shape of the image curve is controllable only up to its "linear" curvature parameters. We present stabilizing control laws for tracking piecewise analytic curves, and propose to track arbitrary curves by approximating them by piecewise "linear" curvature curves. Simulation results are given for these control schemes. Observability of the curve dynamics by using direct measurements from vision sensors as the outputs is studied and an extended Kalman filter is proposed to dynamically estimate the image quantities needed for feedback control from the actual noisy images. Yi Ma 0001, Jana Kosecka, S. Shankar Sastry |
IEEE Trans. Robotics Autom. | 2 |
| 1998 | Motion Recovery from Image Sequences: Discrete Viewpoint vs. Differential Viewpoint
Yi Ma 0001, Jana Kosecka, S. Shankar Sastry |
ECCV (2) | 2 |
| 1998 | A Comparative Study of Vision-Based Lateral Control Strategies for Autonomous Highway DrivingabstractThis paper will present the results of a comparative study of a set of vision-based control strategies that have been applied to the problem of steering an autonomous vehicle along a highway. The aim of this work has been to further our understanding of the characteristics of various control laws that could be applied to this problem with a view to making informed design decisions. The control strategies that we explored include a lead lag control law, a full-state linear controller and input-output linearizing control law. Each of these control strategies was implemented and tested on our experimental vehicle, a Honda Accord LX, both with and without a curvature feedforward component. Jana Kosecka, Robert Blasi, Camillo J. Taylor, Jitendra Malik |
ICRA | 1 |
| 1997 | Generation of conflict resolution manoeuvres for air traffic managementabstractWe explore the use of distributed online motion planning algorithms for multiple mobile agents, in air traffic management systems (ATMS). The work is motivated by current trends in ATMS to move towards decentralized air traffic management, in which the aircraft operate in "free flight" mode instead of following prespecified "sky freeways". Conflict resolution strategies are an integral part of the free flight setting. The purpose of this paper is to obtain a set of manoeuvres to cover all possible conflict scenarios involving multiple agents. A distributed motion planning algorithm based on potential and vortex fields is used. While the algorithm is not always guaranteed to generate flyable trajectories, the obtained trajectories can serve as qualitative prototypes for coordination manoeuvres between multiple aircraft. The actual manoeuvres are generated by approximating these prototypes with trajectories made zip of straight lines and are further verified using hybrid verification techniques. Jana Kosecka, Claire J. Tomlin, George J. Pappas, S. Shankar Sastry |
IROS | 1 |
| 1995 | Cooperative material handling by human and robotic agents: module development and system synthesisabstractPresents a collaborative effort to design and implement a cooperative material handling system by a small team of human and robotic agents in an unstructured indoor environment. The authors' approach makes fundamental use of the human agents' expertise for aspects of task planning, task monitoring, and error recovery. The authors' system is neither fully autonomous nor fully teleoperated. It is designed to make effective use of the human's abilities within the present state of the art of autonomous systems. The authors' robotic agents refer to systems which are each equipped with at least one sensing modality and which possess some capability for self-orientation and/or mobility. The authors' robotic agents are not required to be homogeneous with respect to either capabilities or function. The authors' research stresses both paradigms and testbed experimentation. Theory issues include the requisite coordination principles and techniques which are fundamental to a cooperative multiagent system's basic functioning. The authors have constructed an experimental distributed multiagent-architecture testbed facility. The required modular components of this testbed are currently operational and have been tested individually. The authors' current research focuses on the agents' integration in a scenario for cooperative material handling. Julie A. Adams, Ruzena Bajcsy, Jana Kosecka, Vijay Kumar 0001, Robert Mandelbaum, Max Mintz, Richard P. Paul, Curtis Wang, Yoshio Yamamoto, Xiaoping Yun |
IROS (1) | 3 |
| 1995 | Discrete event modeling of visually guided behaviors
Jana Kosecka, Henrik I. Christensen, Ruzena Bajcsy |
Int. J. Comput. Vis. | 1 |
| 1994 | Application of Discrete Events Systems for Modeling and Controlling Robotic AgentsabstractIn this paper we present a framework for modeling behaviors and tasks for heterogeneous robotic agents. For this purpose we have adopted a formalism from the discrete events systems (DES) theory. We distinguish two kinds of scenarios. In the first one, reactive behaviors of mobile agents directly connect observations with actions. The overall objective is to achieve controllability of the system which is composed from modular components operating in parallel. In the second one, observations are implicitly connected with actions and the objective is to design an observer for manipulatory tasks which would guarantee the task's observability. The use of the DES formalism allows one to describe complex interactions between different components in a systematic fashion and guarantee some control-theoretic properties. We demonstrate our approach by presenting examples of navigation, obstacle avoidance, piercing and picking.> Jana Kosecka, Luca Bogoni |
ICRA | 1 |
| 1993 | Cooperation of visually guided behaviorsabstractThe authors present modeling, analysis, and synthesis of visual behaviors of agents engaged in navigational tasks. They consider situations in which two agents can navigate independently or in cooperation. For the purpose of modeling the behaviors, a formalism is adopted from the discrete events systems (DES) theory that is suitable for investigating control-theoretic issues of a system. The focus is on the identification of elementary behaviors and their composition, leading to more complex behaviors. Two kinds of elementary behaviors are identified: one where observations are directly connected with actions, and one where observations and actions are either received or transmitted. The use of the DES formalism allows synthesis of complex behaviors in a systematic fashion and guarantees their controllability.> Jana Kosecka, Ruzena Bajcsy |
ICCV | 1 |