EDBT 2026 Demo / reviewers in the wild / expert
Eshed Ohn-Bar
dblp:121/0305
· DBLP profile ↗
49ranked-venue papers
18as first author
17since 2021 · last 2025
0000-0002-6750-7987ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 13 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ZeroVO: Visual Odometry with Minimal AssumptionsabstractWe introduce ZeroVO, a novel visual odometry (VO) algorithm that achieves zero-shot generalization across diverse cameras and environments, overcoming limitations in existing methods that depend on predefined or static camera calibration setups. Our approach incorporates three main innovations. First, we design a calibration-free, geometry-aware network structure capable of handling noise in estimated depth and camera parameters. Second, we introduce a language-based prior that infuses semantic information to enhance robust feature extraction and generalization to previously unseen domains. Third, we develop a flexible, semi-supervised training paradigm that iteratively adapts to new scenes using unlabeled data, further boosting the models’ ability to generalize across diverse real-world scenarios. We analyze complex autonomous driving contexts, demonstrating over 30% improvement against prior methods on three standard benchmarks—KITTI, nuScenes, and Argoverse 2—as well as a newly introduced, high-fidelity synthetic dataset derived from Grand Theft Auto (GTA). By not requiring fine-tuning or camera calibration, our work broadens the applicability of VO, providing a versatile solution for real-world deployment at scale. Lei Lai, Zekai Yin, Eshed Ohn-Bar |
CVPR | 3 |
| 2025 | Passing the Driving Knowledge Test
Maolin Wei, Wanzhou Liu, Eshed Ohn-Bar |
ICCV | 3 |
| 2025 | Scalable Offline Metrics for Autonomous DrivingabstractReal-World evaluation of perception-based planning models for robotic systems, such as autonomous vehicles, can be safely and inexpensively conducted offline, ı.e., by computing model prediction error over a pre-collected validation dataset with ground-truth annotations. However, extrapolating from offline model performance to online settings remains a challenge. In these settings, seemingly minor errors can compound and result in test-time infractions or collisions. This relationship is understudied, particularly across diverse closed-loop metrics and complex urban maneuvers. In this work, we revisit this undervalued question in policy evaluation through an extensive set of experiments across diverse conditions and metrics. Based on analysis in simulation, we find an even worse correlation between offline and online settings than reported by prior studies, casting doubts on the validity of current evaluation practices and metrics for driving policies. Next, we bridge the gap between offline and online evaluation. We investigate an offline metric based on epistemic uncertainty, which aims to capture events that are likely to cause errors in closed-loop settings. The resulting metric achieves over 13% improvement in correlation compared to previous offline metrics. We further validate the generalization of our findings beyond the simulation environment in real-world settings, where even greater gains are observed. Animikh Aich, Adwait Kulkarni, Eshed Ohn-Bar |
IROS | 3 |
| 2025 | Navigating the challenges of remotely supporting blind riders in ridesharing
Eshed Ohn-Bar, Ruizhao Zhu, Jimuyang Zhang |
Int. J. Hum. Comput. Stud. | 1 |
| 2024 | Motion Diversification NetworksabstractWe introduce Motion Diversification Networks, a novel framework for learning to generate realistic and diverse 3D human motion. Despite recent advances in deep generative motion modeling, existing models often fail to produce samples that capture the full range of plausible and natural 3D human motion within a given context. The lack of diversity becomes even more apparent in applications where subtle and multi-modal 3D human forecasting is crucial for safety, such as robotics and autonomous driving. Towards more realistic and functional 3D motion models, we highlight limitations in existing generative modeling techniques, particularly in overly simplistic latent code sampling strategies. We then introduce a transformer-based diversification mechanism that learns to effectively guide sampling in the latent space. Our proposed attention-based module queries multiple stochastic samples to flexibly predict a diverse set of latent codes which can be subsequently decoded into motion samples. The proposed framework achieves state-of-the-art diversity and accuracy prediction performance across a range of benchmarks and settings, particularly when used to forecast intricate in-the-wild 3D human motion within complex urban environments. Our models, datasets, and code are available at https://mdncvpr.github.iol. Hee Jae Kim, Eshed Ohn-Bar |
CVPR | 2 |
| 2024 | Uncertainty-Guided Never-Ending Learning to DriveabstractWe present a highly scalable self-training framework for incrementally adapting vision-based end-to-end autonomous driving policies in a semi-supervised manner, i.e., over a continual stream of incoming video data. To facilitate large-scale model training (e.g., open web or unlabeled data), we do not assume access to ground-truth labels and instead estimate pseudo-label policy targets for each video. Our framework comprises three key components: knowledge distillation, a sample purification module, and an exploration and knowledge retention mechanism. First, given sequential image frames, we pseudo-label the data and estimate uncertainty using an ensemble of inverse dynamics models. The uncertainty is used to select the most informative samples to add to an experience replay buffer. We specifically select high-uncertainty pseudo-labels to facilitate the exploration and learning of new and diverse driving skills. However, in contrast to prior work in continual learning that assumes ground-truth labeled samples, the uncertain pseudo-labels can introduce significant noise. Thus, we also pair the exploration with a label refinement module, which makes use of consistency constraints to re-label the noisy exploratory samples and effectively learn from diverse data. Trained as a complete never-ending learning system, we demonstrate state-of-the-art performance on training from domain-changing data as well as millions of images from the open web. Lei Lai, Eshed Ohn-Bar, Sanjay Arora, John Seon Keun Yi |
CVPR | 2 |
| 2024 | Feedback-Guided Autonomous DrivingabstractWhile behavior cloning has recently emerged as a highly successful paradigm for autonomous driving, humans rarely learn to perform complex tasks, such as driving, via imitation or behavior cloning alone. In contrast, learning in humans often involves additional detailed guidance throughout the interactive learning process, i.e., where feedback, often via language, provides detailed information as to which part of their trial was performed incorrectly or suboptimally and why. Motivated by this observation, we introduce an efficient feedback-based framework for improving behavior-cloning-based training of sensorimotor driving agents. Our key insight is to leverage recent advances in Large Language Models (LLMs) to provide corrective fine-grained feedback regarding the underlying reason behind driving prediction failures. Moreover, our introduced network architecture is efficient, enabling the first sensorimotor end-to-end training and evaluation of LLM-based driving models. The resulting agent achieves state-of-the-art performance in open-loop evaluation on nuScenes, outperforming prior state-of-the-art by over 8.1% and 57.1% in accuracy and collision rate, respectively. In CARLA, our camera-based agent improves by 16.6% in driving score over prior LIDAR-based approaches. Jimuyang Zhang, Zanming Huang, Arijit Ray, Eshed Ohn-Bar |
CVPR | 4 |
| 2024 | Neural Volumetric World Models for Autonomous Driving
Zanming Huang, Jimuyang Zhang, Eshed Ohn-Bar |
ECCV (17) | 3 |
| 2024 | Unified Local-Cloud Decision-Making via Reinforcement Learning
Kathakoli Sengupta, Zhongkai Shangguan, Sandesh Bharadwaj, Sanjay Arora, Eshed Ohn-Bar, Renato Mancuso 0001 |
ECCV (41) | 5 |
| 2024 | Text to Blind MotionabstractPeople who are blind perceive the world differently than those who are sighted, which can result in distinct motion characteristics. For instance, when crossing at an intersection, blind individuals may have different patterns of movement, such as veering more from a straight path or using touch-based exploration around curbs and obstacles. These behaviors may appear less predictable to motion models embedded in technologies such as autonomous vehicles. Yet, the ability of 3D motion models to capture such behavior has not been previously studied, as existing datasets for 3D human motion currently lack diversity and are biased toward people who are sighted. In this work, we introduce BlindWays, the first multimodal motion benchmark for pedestrians who are blind. We collect 3D motion data using wearable sensors with 11 blind participants navigating eight different routes in a real-world urban setting. Additionally, we provide rich textual descriptions that capture the distinctive movement characteristics of blind pedestrians and their interactions with both the navigation aid (e.g., a white cane or a guide dog) and the environment. We benchmark state-of-the-art 3D human prediction models, finding poor performance with off-the-shelf and pre-training-based methods for our novel task. To contribute toward safer and more reliable systems that can seamlessly reason over diverse human movements in their environments, our text-and-motion benchmark is available at https://blindways.github.io/. Hee Jae Kim, Kathakoli Sengupta, Masaki Kuribayashi, Hernisa Kacorri, Eshed Ohn-Bar |
NeurIPS | 5 |
| 2024 | Scalable Early Childhood Reading Performance PredictionabstractModels for student reading performance can empower educators and institutions to proactively identify at-risk students, thereby enabling early and tailored instructional interventions. However, there are no suitable publicly available educational datasets for modeling and predicting future reading performance. In this work, we introduce the Enhanced Core Reading Instruction (ECRI) dataset, a novel large-scale longitudinal tabular dataset collected across 44 schools with 6,916 students and 172 teachers. We leverage the dataset to empirically evaluate the ability of state-of-the-art machine learning models to recognize early childhood educational patterns in multivariate and partial measurements. Specifically, we demonstrate a simple self-supervised strategy in which a Multi-Layer Perception (MLP) network is pre-trained over masked inputs to outperform several strong baselines while generalizing over diverse educational settings. To facilitate future developments in precise modeling and responsible use of models for individualized and early intervention strategies, our data and code are available at https://ecri-data.github.io/. Zhongkai Shangguan, Zanming Huang, Eshed Ohn-Bar, Ola Ozernov-Palchik, Derek Kosty, Michael Stoolmiller, Hank Fien |
NeurIPS | 3 |
| 2023 | Coaching a Teachable StudentabstractWe propose a novel knowledge distillation framework for effectively teaching a sensorimotor student agent to drive from the supervision of a privileged teacher agent. Current distillation for sensorimotor agents methods tend to result in suboptimal learned driving behavior by the student, which we hypothesize is due to inherent differences between the input, modeling capacity, and optimization processes of the two agents. We develop a novel distillation scheme that can address these limitations and close the gap between the sensorimotor agent and its privileged teacher. Our key insight is to design a student which learns to align their input features with the teacher's privileged Bird's Eye View (BEV) space. The student then can benefit from direct supervision by the teacher over the internal representation learning. To scaffold the difficult sensorimotor learning task, the student model is optimized via a student-paced coaching mechanism with various auxiliary supervision. We further propose a high-capacity imitation learned privileged agent that surpasses prior privileged agents in CARLA and ensures the student learns safe driving behavior. Our proposed sensorimotor agent results in a robust image-based behavior cloning agent in CARLA, improving over current models by over 20.6% in driving score without requiring LiDAR, historical observations, ensemble of models, on-policy data aggregation or reinforcement learning. Jimuyang Zhang, Zanming Huang, Eshed Ohn-Bar |
CVPR | 3 |
| 2023 | XVO: Generalized Visual Odometry via Cross-Modal Self-TrainingabstractWe propose XVO, a semi-supervised learning method for training generalized monocular Visual Odometry (VO) models with robust off-the-self operation across diverse datasets and settings. In contrast to standard monocular VO approaches which often study a known calibration within a single dataset, XVO efficiently learns to recover relative pose with real-world scale from visual scene semantics, i.e., without relying on any known camera parameters. We optimize the motion estimation model via self-training from large amounts of unconstrained and heterogeneous dash camera videos available on YouTube. Our key contribution is twofold. First, we empirically demonstrate the benefits of semi-supervised training for learning a general-purpose direct VO regression network. Second, we demonstrate multi-modal supervision, including segmentation, flow, depth, and audio auxiliary prediction tasks, to facilitate generalized representations for the VO task. Specifically, we find audio prediction task to significantly enhance the semi-supervised learning process while alleviating noisy pseudo-labels, particularly in highly dynamic and out-of-domain video data. Our proposed teacher network achieves state-of-the-art performance on the commonly used KITTI benchmark despite no multi-frame optimization or knowledge of camera parameters. Combined with the proposed semi-supervised step, XVO demonstrates off-the-shelf knowledge transfer across diverse conditions on KITTI, nuScenes, and Argoverse without fine-tuning. Lei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-Bar |
ICCV | 4 |
| 2022 | SelfD: Self-Learning Large-Scale Driving Policies From the WebabstractEffectively utilizing the vast amounts of ego-centric navigation data that is freely available on the internet can advance generalized intelligent systems, i.e., to robustly scale across perspectives, platforms, environmental conditions, scenarios, and geographical locations. However, it is difficult to directly leverage such large amounts of unlabeled and highly diverse datafor complex 3D reasoning and planning tasks. Consequently, researchers have primarily focused on its use for various auxiliary pixel- and image-level computer vision tasks that do not consider an ultimate navigational objective. In this work, we introduce SelfD, a framework for learning scalable driving by utilizing large amounts of online monocular images. Our key idea is to leverage iterative semi-supervised training when learning imitative agents from unlabeled data. To handle unconstrained viewpoints, scenes, and camera parameters, we train an image-based model that directly learns to plan in the Bird's Eye View (BEV) space. Next, we use unla-beled data to augment the decision-making knowledge and robustness of an initially trained model via self-training. In particular, we propose a pseudo-labeling step which enables making full use of highly diverse demonstration data through “hypothetical” planning-based data augmentation. We employ a large dataset of publicly available YouTube videos to train SelfD and comprehensively analyze its generalization benefits across challenging navigation scenarios. Without requiring any additional data collection or annotation efforts, SelfD demonstrates consistent improvements (by up to 24%) in driving performance evaluation on nuScenes, Argoverse, Waymo, and CARLA. Jimuyang Zhang, Ruizhao Zhu, Eshed Ohn-Bar |
CVPR | 3 |
| 2022 | ASSISTER: Assistive Navigation via Conditional Instruction Generation
Zanming Huang, Zhongkai Shangguan, Jimuyang Zhang, Gilad Bar, Matthew Boyd, Eshed Ohn-Bar |
ECCV (36) | 6 |
| 2021 | Learning by WatchingabstractWhen in a new situation or geographical location, human drivers have an extraordinary ability to watch others and learn maneuvers that they themselves may have never performed. In contrast, existing techniques for learning to drive preclude such a possibility as they assume direct access to an instrumented ego-vehicle with fully known observations and expert driver actions. However, such measurements cannot be directly accessed for the non-ego vehicles when learning by watching others. Therefore, in an application where data is regarded as a highly valuable asset, current approaches completely discard the vast portion of the training data that can be potentially obtained through indirect observation of surrounding vehicles. Motivated by this key insight, we propose the Learning by Watching (LbW) framework which enables learning a driving policy without requiring full knowledge of neither the state nor expert actions. To increase its data, i.e., with new perspectives and maneuvers, LbW makes use of the demonstrations of other vehicles in a given scene by (1) transforming the egovehicle’s observations to their points of view, and (2) inferring their expert actions. Our LbW agent learns more robust driving policies while enabling data-efficient learning, including quick adaptation of the policy to rare and novel scenarios. In particular, LbW drives robustly even with a fraction of available driving data required by existing methods, achieving an average success rate of 92% on the original CARLA benchmark with only 30 minutes of total driving data and 82% with only 10 minutes. Jimuyang Zhang, Eshed Ohn-Bar |
CVPR | 2 |
| 2021 | X-World: Accessibility, Vision, and Autonomy MeetabstractAn important issue facing vision-based intelligent systems today is the lack of accessibility-aware development. A main reason for this issue is the absence of any large-scale, standardized vision benchmarks that incorporate relevant tasks and scenarios related to people with disabilities. This lack of representation hinders even preliminary analysis with respect to underlying pose, appearance, and occlusion characteristics of diverse pedestrians. What is the impact of significant occlusion from a wheelchair on instance segmentation quality? How can interaction with mobility aids, e.g., a long and narrow walking cane, be recognized robustly? To begin addressing such questions, we introduce X-World, an accessibility-centered development environment for vision-based autonomous systems. We tackle inherent data scarcity by leveraging a simulation environment to spawn dynamic agents with various mobility aids. The simulation supports generation of ample amounts of finely annotated, multi-modal data in a safe, cheap, and privacy-preserving manner. Our analysis highlights novel challenges introduced by our benchmark and tasks, as well as numerous opportunities for future developments. We further broaden our analysis using a complementary real-world evaluation benchmark of in-situ navigation by pedestrians with disabilities. Our contributions provide an initial step towards widespread deployment of vision-based agents that can perceive and model the interaction needs of diverse people with disabilities. Jimuyang Zhang, Minglan Zheng, Matthew Boyd, Eshed Ohn-Bar |
ICCV | 4 |
| 2020 | Learning Situational DrivingabstractHuman drivers have a remarkable ability to drive in diverse visual conditions and situations, e.g., from maneuvering in rainy, limited visibility conditions with no lane markings to turning in a busy intersection while yielding to pedestrians. In contrast, we find that state-of-the-art sensorimotor driving models struggle when encountering diverse settings with varying relationships between observation and action. To generalize when making decisions across diverse conditions, humans leverage multiple types of situation-specific reasoning and learning strategies. Motivated by this observation, we develop a framework for learning a situational driving policy that effectively captures reasoning under varying types of scenarios. Our key idea is to learn a mixture model with a set of policies that can capture multiple driving modes. We first optimize the mixture model through behavior cloning and show it to result in significant gains in terms of driving performance in diverse conditions. We then refine the model by directly optimizing for the driving task itself, i.e., supervised with the navigation task reward. Our method is more scalable than methods assuming access to privileged information, e.g., perception labels, as it only assumes demonstration and reward-based supervision. We achieve over 98% success rate on the CARLA driving benchmark as well as state-of-the-art performance on a newly introduced generalization benchmark. Eshed Ohn-Bar, Aditya Prakash 0001, Aseem Behl, Kashyap Chitta, Andreas Geiger 0001 |
CVPR | 1 |
| 2020 | Exploring Data Aggregation in Policy Learning for Vision-Based Urban Autonomous DrivingabstractData aggregation techniques can significantly improve vision-based policy learning within a training environment, e.g., learning to drive in a specific simulation condition. However, as on-policy data is sequentially sampled and added in an iterative manner, the policy can specialize and overfit to the training conditions. For real-world applications, it is useful for the learned policy to generalize to novel scenarios that differ from the training conditions. To improve policy learning while maintaining robustness when training end-to-end driving policies, we perform an extensive analysis of data aggregation techniques in the CARLA environment. We demonstrate how the majority of them have poor generalization performance, and develop a novel approach with empirically better generalization performance compared to existing techniques. Our two key ideas are (1) to sample critical states from the collected on-policy data based on the utility they provide to the learned policy in terms of driving behavior, and (2) to incorporate a replay buffer which progressively focuses on the high uncertainty regions of the policy's state distribution. We evaluate the proposed approach on the CARLA NoCrash benchmark, focusing on the most challenging driving scenarios with dense pedestrian and vehicle traffic. Our approach improves driving success rate by 16% over state-of-the-art, achieving 87% of the expert performance while also reducing the collision rate by an order of magnitude without the use of any additional modality, auxiliary tasks, architectural modifications or reward from the environment. Aditya Prakash 0001, Aseem Behl, Eshed Ohn-Bar, Kashyap Chitta, Andreas Geiger 0001 |
CVPR | 3 |
| 2020 | Label Efficient Visual Abstractions for Autonomous DrivingabstractIt is well known that semantic segmentation can be used as an effective intermediate representation for learning driving policies. However, the task of street scene semantic segmentation requires expensive annotations. Furthermore, segmentation algorithms are often trained irrespective of the actual driving task, using auxiliary image-space loss functions which are not guaranteed to maximize driving metrics such as safety or distance traveled per intervention. In this work, we seek to quantify the impact of reducing segmentation annotation costs on learned behavior cloning agents. We analyze several segmentation-based intermediate representations. We use these visual abstractions to systematically study the trade-off between annotation efficiency and driving performance, i.e., the types of classes labeled, the number of image samples used to learn the visual abstraction model, and their granularity (e.g., object masks vs. 2D bounding boxes). Our analysis uncovers several practical insights into how segmentation-based visual abstractions can be exploited in a more label efficient manner. Surprisingly, we find that state-of-the-art driving performance can be achieved with orders of magnitude reduction in annotation cost. Beyond label efficiency, we find several additional training benefits when leveraging visual abstractions, such as a significant reduction in the variance of the learned policy when compared to state-of-the-art end-to-end driving models. Aseem Behl, Kashyap Chitta, Aditya Prakash 0001, Eshed Ohn-Bar, Andreas Geiger 0001 |
IROS | 4 |
| 2020 | Virtual navigation for blind people: Transferring route knowledge to the real-World
João Guerreiro 0002, Daisuke Sato 0001, Dragan Ahmetovic, Eshed Ohn-Bar, Kris Makoto Kitani, Chieko Asakawa |
Int. J. Hum. Comput. Stud. | 4 |
| 2020 | Learning Context-dependent Personal Preferences for Adaptive RecommendationabstractWe propose two online-learning algorithms for modeling the personal preferences of users of interactive systems. The proposed algorithms leverage user feedback to estimate user behavior and provide personalized adaptive recommendation for supporting context-dependent decision-making. We formulate preference modeling as online prediction algorithms over a set of learned policies, i.e., policies generated via supervised learning with interaction and context data collected from previous users. The algorithms then adapt to a target user by learning the policy that best predicts that user’s behavior and preferences. We also generalize the proposed algorithms for a more challenging learning case in which they are restricted to a limited number of trained policies at each timestep, i.e., for mobile settings with limited resources. While the proposed algorithms are kept general for use in a variety of domains, we developed an image-filter-selection application. We used this application to demonstrate how the proposed algorithms can quickly learn to match the current user’s selections. Based on these evaluations, we show that (1) the proposed algorithms exhibit better prediction accuracy compared to traditional supervised learning and bandit algorithms, (2) our algorithms are robust under challenging limited prediction settings in which a smaller number of expert policies is assumed. Finally, we conducted a user study to demonstrate how presenting users with the prediction results of our algorithms significantly improves the efficiency of the overall interaction experience. Keita Higuchi, Hiroki Tsuchida, Eshed Ohn-Bar, Yoichi Sato 0001, Kris Makoto Kitani |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2019 | A-EXP4: Online Social Policy Learning for Adaptive Robot-Pedestrian InteractionabstractWe study self-supervised adaptation of a robot's policy for social interaction, i.e., a policy for active communication with surrounding pedestrians through audio or visual signals. Inspired by the observation that humans continually adapt their behavior when interacting under varying social context, we propose Adaptive EXP4 (A-EXP4), a novel online learning algorithm for adapting the robot-pedestrian interaction policy. To address limitations of bandit algorithms in adaptation to unseen and highly dynamic scenarios, we employ a mixture model over the policy parameter space. Specifically, a Dirichlet Process Gaussian Mixture Model (DPMM) is used to cluster the parameters of sampled policies and maintain a mixture model over the clusters, hence effectively discovering policies that are suitable to the current environmental context in an unsupervised manner. Our simulated and real-world experiments demonstrate the feasibility of A-EXP4 in accommodating interaction with different types of pedestrians while jointly minimizing social disruption through the adaptation process. While the A-EXP4 formulation is kept general for application in a variety of domains requiring continual adaptation of a robot's policy, we specifically evaluate the performance of our algorithm using a suitcase-inspired assistive robotic platform. In this concrete assistive scenario, the algorithm observes how audio signals produced by the navigational system affect the behavior of pedestrians and adapts accordingly. Consequently, we find A-EXP4 to effectively adapt the interaction policy for gently clearing a navigation path in crowded settings, resulting in significant reduction in empirical regret compared to the EXP4 baseline. Pengju Jin, Eshed Ohn-Bar, Kris Makoto Kitani, Chieko Asakawa |
IROS | 2 |
| 2019 | Forecasting Time-to-Collision from Monocular Video: Feasibility, Dataset, and ChallengesabstractWe explore the possibility of using a single monocular camera to forecast the time to collision between a suitcase-shaped robot being pushed by its user and other nearby pedestrians. We develop a purely image-based deep learning approach that directly estimates the time to collision without the need of relying on explicit geometric depth estimates or velocity information to predict future collisions. While previous work has focused on detecting immediate collision in the context of navigating Unmanned Aerial Vehicles, the detection was limited to a binary variable (i.e., collision or no collision). We propose a more fine-grained approach to collision forecasting by predicting the exact time to collision in terms of milliseconds, which is more helpful for collision avoidance in the context of dynamic path planning. To evaluate our method, we have collected a novel dataset of over 13,000 indoor video segments each showing a trajectory of at least one person ending in a close proximity (a near collision) with the camera mounted on a mobile suitcase-shaped platform. Using this dataset, we do extensive experimentation on different temporal windows as input using an exhaustive list of state-of-the-art convolutional neural networks (CNNs). Our results show that our proposed multi-stream CNN is the best model for predicting time to near-collision. The average prediction error of our time to near-collision is 0.75 seconds across the test videos. The project webpage can be found at https://aashi7.github.io/NearCollision.html. Aashi Manglik, Xinshuo Weng, Eshed Ohn-Bar, Kris Makoto Kitani |
IROS | 3 |
| 2018 | Environmental Factors in Indoor Navigation Based on Real-World Trajectories of Blind UsersabstractIndoor localization technologies can enhance quality of life for blind people by enabling them to independently explore and navigate indoor environments. Researchers typically evaluate their systems in terms of localization accuracy and user behavior along planned routes. We propose two measures of path-following behavior: deviation from optimal route and trajectory variability. Through regression analysis of real-world trajectories from blind users, we identify relationships between a) these measures and b) elements of the environment, route characteristics, localization error, and instructional cues that users receive. Our results provide insights into path-following behavior for turn-by-turn indoor navigation and have implications for the design of future interactions. Moreover, our findings highlight the importance of reporting these environmental factors and route properties in similar studies. We present automated and scalable methods for their calculation and to encourage their reporting for better interpretation and comparison of results across future studies. Hernisa Kacorri, Eshed Ohn-Bar, Kris Makoto Kitani, Chieko Asakawa |
CHI | 2 |
| 2018 | Modeling Expertise in Assistive Navigation Interfaces for Blind PeopleabstractEvaluating the impact of expertise and route knowledge on task performance can guide the design of intelligent and adaptive navigation interfaces. Expertise has been relatively unexplored in the context of assistive indoor navigation interfaces for blind people. To quantify the complex relationship between the user»s walking patterns, route learning, and adaptation to the interface, we conducted a study with 8 blind participants. The participants repeated a set of navigation tasks while using a smartphone-based turn-by-turn navigation guidance app. The results demonstrate the gradual evolution of user skill and knowledge throughout the route repetitions, significantly impacting the task completion time. In addition to the exploratory analysis, we take a step towards tailoring the navigation interface to the user»s needs by proposing a personalized recurrent neural network-based behavior model for expertise level classification. Eshed Ohn-Bar, João Guerreiro 0002, Dragan Ahmetovic, Kris Makoto Kitani, Chieko Asakawa |
IUI | 1 |
| 2018 | SmartPartNet: Part-Informed Person Detection for Body-Worn SmartphonesabstractWe are interested in the development of image-based person detection algorithms for wearable computing using commodity smartphones. We focus on the use of smartphones as a wearable device because it is a practical means of augmenting human sensing for applications such as navigation for the blind or assisting social interaction. We identify two unique features of developing a vision-based person detector for body-worn smartphones: (1) the detector must take into account the strong bias in the size of people in the images taken with a wearable device and (2) the detector must consider the low image quality due to dim lighting and rapid ego-motion which leads to motion blur. In order to account for the unique distribution over the visibility of body parts when using a wearable camera, we propose a part-based person detector specialized for chestmounted smartphones. We perform extensive ablative analysis on the usefulness of part information, providing several insights regarding the design of the optimal person detector across different application domains. To account for the frequent occurrence of motion blur in our target domain, we introduce a data augmentation technique to generate synthetic motion-blurred images during training. In addition to addressing the aforementioned features, the final detector must also run in real-time using only smartphone resources. We leverage recent progress in deep neural networks for mobile devices and show that our proposed person detector, SmartPartNet, obtains performance similar to state-of-the-art pedestrian detection networks, while being 3X smaller and 5X faster. Heng Yu 0005, Eshed Ohn-Bar, Donghyun Yoo, Kris Makoto Kitani |
WACV | 2 |
| 2017 | Multi-scale volumes for deep object detection and localization
Eshed Ohn-Bar, Mohan M. Trivedi |
Pattern Recognit. | 1 |
| 2017 | Are all objects equal? Deep spatio-temporal importance prediction in driving videos
Eshed Ohn-Bar, Mohan M. Trivedi |
Pattern Recognit. | 1 |
| 2016 | Preparatory coordination of head, eyes and hands: Experimental study at intersectionsabstractDrivers use some combination of head, eye and hand movements to perform varying number of tasks from driving related to non-driving secondary tasks. Furthermore, the combinations may vary depending on the task performed. It is important to model and understand these variations in order to build predictive systems, explore driving styles, detect activities, etc. This study, therefore, introduces a framework to model the spatio-temporal movements of head, eyes and hands given naturalistic driving data of looking-in at the driver for any events or tasks performed of interest. As a use case, we explore the temporal coordination of the modalities on data of drivers executing maneuvers at stop-controlled intersections; the maneuvers executed are go straight, turn left and turn right. In sequentially increasing time windows, by training classifiers which have the ability to provide discriminative quality of its input variable, the experimental study at intersections shows which type of, when and how long distinguishable preparatory movements occur in the range of a few milliseconds to a few seconds. Sujitha Martin, Akshay Rangesh, Eshed Ohn-Bar, Mohan M. Trivedi |
ICPR | 3 |
| 2016 | Detection and localization with multi-scale modelsabstractObject detection and localization in images involve a multi-scale reasoning process. First, responses of object detectors are known to vary with image scale. Second, contextual relationships on a part-level, object-level, and scene-level appear at different scales of the image. This paper studies efficient modeling of these two components by training multi-scale template models. The input to the proposed algorithm involves image features computed at varying image scales, hence operating on volumes in the feature pyramid. The approach generalizes single-scale, local-region detection approaches (e.g. sliding window or region proposals), jointly learning detection and localization cues. Extending the common single-scale detection to a multi-scale volume allows learning scale-specific models as well as analyzing the importance of contextual information at different scales. Experimental analysis on the PASCAL VOC dataset shows the method to considerably improve both detection and localization performance for different type of features, histogram of oriented gradients and deep convolutional neural network features. Eshed Ohn-Bar, Mohan M. Trivedi |
ICPR | 1 |
| 2016 | To boost or not to boost? On the limits of boosted trees for object detectionabstractWe aim to study the modeling limitations of the commonly employed boosted decision trees classifier. Inspired by the success of large, data-hungry visual recognition models (e.g. deep convolutional neural networks), this paper focuses on the relationship between modeling capacity of the weak learners, dataset size, and dataset properties. A set of novel experiments on the Caltech Pedestrian Detection benchmark results in the best known performance among non-CNN techniques while operating at fast run-time speed. Furthermore, the performance is on par with deep architectures (9.71% log-average miss rate), while using only HOG+LUV channels as features. The conclusions from this study are shown to generalize over different object detection domains as demonstrated on the FDDB face detection benchmark (93.37% accuracy). Despite the impressive performance, this study reveals the limited modeling capacity of the common boosted trees model, motivating a need for architectural changes in order to compete with multi-level and very deep architectures. Eshed Ohn-Bar, Mohan M. Trivedi |
ICPR | 1 |
| 2016 | What makes an on-road object important?abstractHuman drivers continuously attend to important scene elements in order to safely and smoothly navigate in intricate environments and under uncertainty. This paper develops a human-centric framework for object recognition by analyzing a notion of object importance, as measured in a spatio-temporal context of driving a vehicle. Given a video, a main research question in this paper is - which of the surrounding agents are most important? The answer inherently requires complex reasoning over the current driving task, object properties, scene context, intent, and possible future actions. Therefore, we find that various spatio-temporal cues are relevant for the importance classification task. Furthermore, we demonstrate the usefulness of the importance annotations in evaluating vision algorithms (specifically, for the task of object detection) in an application where trust in automation is imperative and errors are costly. Finally, we show that importance-guided training of object detection models results in improved detection performance of surrounding objects of higher importance. Hence, such models may be better suited for use in representing safety-critical situations, predicting surrounding agents' intentions, and in human-robot interactivity. The dataset and code will be made publicly available. Eshed Ohn-Bar, Mohan M. Trivedi |
ICPR | 1 |
| 2016 | The rhythms of head, eyes and hands at intersectionsabstractIn this paper, we study the complex coordination of head, eyes and hands as the driver approaches a stop-controlled intersection. The proposed framework is made up of three major parts. The first part is the naturalistic driving dataset collection: synchronized capture of sensors looking-in and looking-out, multiple drivers driving in urban environment, and segmenting events at stop-controlled intersections. The second part is extracting reliable features from purely vision sensors looking in at the driver: eye movements, head pose and hand location respective to the wheel. The third part is in the design of appropriate temporal features for capturing coordination. A random forest algorithm is employed for studying relevance and understanding the temporal evolution of head, eye, and hand cues. Using 24 different events (from 5 drivers resulting in ~ 12200 frames analyzed) of three different maneuvers at stop-controlled intersections, we found that preparatory motions range in the order of a few seconds to a few milliseconds, depending on the modality (i.e. eyes, head, hand), before the event occurs. Sujitha Martin, Akshay Rangesh, Eshed Ohn-Bar, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 3 |
| 2016 | Looking at Pedestrians at Different Scales: A Multiresolution Approach and EvaluationsabstractTypically, in a detector framework, the model size is fixed at the size of the smallest object to be detected, and larger objects are detected by scaling the input image. The information lost due to scaling could be vital for accurately detecting large objects, which is an essential task for vision-based driver-assistance systems. To this end, we evaluate a multiresolution detector framework by training models at different sizes and demonstrate its effectiveness on a state-of-the-art pedestrian detector. Our comprehensive evaluation demonstrates meaningful improvement in detector performance. On the KITTI dataset under moderate difficulty settings, we achieve a 6% increase in the detector's average precision over the baseline single-resolution result on the KITTI benchmark. Further insights into the detector's improvements are provided using a fine-grained analysis of the detector's performance at various threshold settings. Rakesh Nattoji Rajaram, Eshed Ohn-Bar, Mohan M. Trivedi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2016 | Long-Term Multi-Cue Tracking of Hands in VehiclesabstractHands are a very important cue for understanding and analyzing driver activity and human activity, in general. Vision-based hand detection and tracking involve major challenges, such as attaining robustness to inconsistencies in lighting and scale, background clutter, object occlusion/disappearance and the large variability in hand shape, size, color, and structure. In this paper, we introduce a novel framework suitable for tracking multiple hands online. Assigning tracks to these detections is modeled as a bipartite matching problem with an objective of minimizing the total cost. Both motion and appearance cues are integrated in order to gain robustness to occlusion, fast movement, and interacting hands. Additionally, we study the utility of a left versus right hand classifier to disambiguate hand tracks and reduce ID switches. The proposed tracker shows promise on an extensive, naturalistic, and publicly available driving (VIVA Challenge) data set, by tracking both hands of the driver and the passenger effectively. Akshay Rangesh, Eshed Ohn-Bar, Mohan M. Trivedi |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Can appearance patterns improve pedestrian detection?abstractThis paper studies the usefulness of appearance patterns for the challenging task of pedestrian detection. Despite appearance specific models being common in rigid object detection, the technique is still little understood for pedestrians. Three main approaches for reasoning over orientation, occlusion, and visual cues in obtaining the appearance patterns are compared. This work demonstrates that large gains in detection performance (up to 17 AP points on the challenging KITTI dataset) can be made using a state-of-the-art pedestrian detector. Eshed Ohn-Bar, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 1 |
| 2015 | A comparative study of color and depth features for hand gesture recognition in naturalistic driving settingsabstractWe are concerned with investigating efficient video representations for the purpose of hand gesture recognition in settings of naturalistic driving. In order to provide a common experimental setup for previously proposed space-time features, we study a color and depth naturalistic hand gesture benchmark. The dataset allows for evaluation of descriptors under settings of common self-occlusion and large illumination variation. A collection of simple and quick to extract spatio-temporal cues requiring no codebook encoding are proposed. Their effectiveness is validated on our dataset, as well as on the Cambridge hand gesture dataset, improving state-of-the-art. Finally, fusion of the modalities and various cues is studied. Eshed Ohn-Bar, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 1 |
| 2015 | On surveillance for safety critical events: In-vehicle video networks for predictive driver assistance systems
Eshed Ohn-Bar, Ashish Tawari, Sujitha Martin, Mohan M. Trivedi |
Comput. Vis. Image Underst. | 1 |
| 2015 | Learning to Detect Vehicles by Clustering Appearance PatternsabstractThis paper studies efficient means in dealing with intracategory diversity in object detection. Strategies for occlusion and orientation handling are explored by learning an ensemble of detection models from visual and geometrical clusters of object instances. An AdaBoost detection scheme is employed with pixel lookup features for fast detection. The analysis provides insight into the design of a robust vehicle detection system, showing promise in terms of detection performance and orientation estimation accuracy. Eshed Ohn-Bar, Mohan M. Trivedi |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2014 | Head, Eye, and Hand Patterns for Driver Activity RecognitionabstractIn this paper, a multiview, multimodal vision framework is proposed in order to characterize driver activity based on head, eye, and hand cues. Leveraging the three types of cues allows for a richer description of the driver's state and for improved activity detection performance. First, regions of interest are extracted from two videos, one observing the driver's hands and one the driver's head. Next, hand location hypotheses are generated and integrated with a head pose and facial landmark module in order to classify driver activity into three states: wheel region interaction with two hands on the wheel, gear region activity, or instrument cluster region activity. The method is evaluated on a video dataset captured in on-road settings. Eshed Ohn-Bar, Sujitha Martin, Ashish Tawari, Mohan M. Trivedi |
ICPR | 1 |
| 2014 | Go with the Flow: Improving Multi-view Vehicle Detection with Motion CuesabstractAs vehicles travel through a scene, changes in aspect ratio and appearance as observed from a camera (or an array of cameras) make vehicle detection a difficult computer vision problem. Rather than relying solely on appearance cues, we propose a framework for detecting vehicles and eliminating false positives by utilizing the motion cues in the scene in addition to the appearance cues. As a case study, we focus on overtaking vehicle detection in a freeway setting from forward and rear views of the ego-vehicle. The proposed integration occurs in two steps. First, motion-based vehicle detection is performed using optical flow. Taking advantage of epipolar constraints, salient motion vectors are extracted and clustered using spectral clustering to form bounding boxes of vehicle candidates. Post-processing and outlier removal further refine the detections. Second, the motion-based detections are then combined with the output of an appearance-based vehicle detector to reduce false positives and produce the final vehicle detections. Alfredo Ramirez, Eshed Ohn-Bar, Mohan M. Trivedi |
ICPR | 2 |
| 2014 | Understanding head and hand activities and coordination in naturalistic driving videosabstractIn this work, we propose a vision-based analysis framework for recognizing in-vehicle activities such as interactions with the steering wheel, the instrument cluster and the gear. The framework leverages two views for activity analysis, a camera looking at the driver's hand and another looking at the driver's head. The techniques proposed can be used by researchers in order to extract ‘mid-level’ information from video, which is information that represents some semantic understanding of the scene but may still require an expert in order to distinguish difficult cases or leverage the cues to perform drive analysis. Unlike such information, ‘low-level’ video is large in quantity and can't be used unless processed entirely by an expert. This work can apply to minimizing manual labor so that researchers may better benefit from the accessibility of the data and provide them with the ability to perform larger-scaled studies. Sujitha Martin, Eshed Ohn-Bar, Ashish Tawari, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 2 |
| 2014 | Predicting driver maneuvers by learning holistic featuresabstractIn this work, we propose a framework for the recognition and prediction of driver maneuvers by considering holistic cues. With an array of sensors, driver's head, hand, and foot gestures are being captured in a synchronized manner together with lane, surrounding agents, and vehicle parameters. An emphasis is put on real-time algorithms. The cues are processed and fused using a latent-dynamic discriminative framework. As a case study, driver activity recognition and prediction in overtaking situations is performed using a naturalistic, on-road dataset. A consequence of this work would be in development of more effective driver analysis and assistance systems. Eshed Ohn-Bar, Ashish Tawari, Sujitha Martin, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 1 |
| 2014 | Integrating motion and appearance for overtaking vehicle detectionabstractThe dynamic appearance of vehicles as they enter and exit a scene makes vehicle detection a difficult and complicated problem. Appearance based detectors generally provide good results when vehicles are in clear view, but have trouble in the scenes edges due to changes in the vehicles aspect ratio and partial occlusions. To compensate for some of these deficiencies, we propose incorporating motion cues from the scene. In this work, we focus on a overtaking vehicle detection in a freeway setting with front and rear facing monocular cameras. Motion cues are extracted from the scene, and leveraging the epipolar geometry of the monocular setup, motion compensation is performed. Spectral clustering is used to group similar motion vectors together, and after post-processing, vehicle detections candidates are produced. Finally, these candidates are combined with an appearance detector to remove any false positives, outputting the detections as a vehicle travels through the scene. Alfredo Ramirez, Eshed Ohn-Bar, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 2 |
| 2014 | Hand Gesture Recognition in Real Time for Automotive Interfaces: A Multimodal Vision-Based Approach and EvaluationsabstractIn this paper, we develop a vision-based system that employs a combined RGB and depth descriptor to classify hand gestures. The method is studied for a human-machine interface application in the car. Two interconnected modules are employed: one that detects a hand in the region of interaction and performs user classification, and another that performs gesture recognition. The feasibility of the system is demonstrated using a challenging RGBD hand gesture data set collected under settings of common illumination variation and occlusion. Eshed Ohn-Bar, Mohan M. Trivedi |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2013 | Partially occluded vehicle recognition and tracking in 3DabstractAbstract — Vehicle detection is a key problem in computer vision, with applications in driver assistance and active safety. A challenging aspect of the problem is the common occlusion of vehicles in the scene. In this paper, we present a visionbased system for vehicle localization and tracking for detecting partially visible vehicles. Consequently, vehicles are localized more reliably and tracked for longer periods of time. The proposed system detects vehicles using an active-learning based monocular vision approach and motion (optical flow) cues. A calibrated stereo rig is utilized to acquire a depth map, and consequently the real-world coordinates of each detected vehicle. Tracking is performed using a Kalman filter. The tracking is formulated to integrate stereo-monocular information. We demonstrate the effectiveness of the proposed system on a multilane highway dataset containing instances of vehicles with relative motion to the ego-vehicle. I. Eshed Ohn-Bar, Sayanan Sivaraman, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 1 |
| 2013 | In-vehicle hand activity recognition using integration of regionsabstractIn this paper, we focus on the analysis of naturalistic driver behavior using hand activity. To that end, a dataset of color and depth images under varying operating modes and illumination settings was collected. The proposed framework provides a robust solution for localizing the hands by partitioning visible and depth images into disjoint sub-regions which may be of interest for studying the state of the driver: wheel, lap, hand rest, gear, and infotainment region. Different feature extraction methods are proposed and thoroughly studied in terms of speed and performance for each of the five regions. A model for hand presence is learned for each region separately, and these are integrated using a second-stage classifier. As the appearance of hands varies among regions and the hands can only be found in a subset of the regions chosen, the technique leverages information and confidence from multiple regions to produce hand activity classification. Eshed Ohn-Bar, Mohan M. Trivedi |
Intelligent Vehicles Symposium | 1 |
| 2012 | Hand gesture-based visual user interface for infotainmentabstractWe present a real-time vision-based system that discriminates hand gestures performed by in-vehicle front-row seat occupants for accessing the infotainment system. The hand gesture-based visual user interface may be more natural and intuitive to the user than the current tactile interaction interface. Consequently, it may encourage a gaze-free interaction, which can alleviate driver distraction without limiting the user's infotainment experience. The system uses visible and depth images of the dashboard and center-console area in the vehicle. The first step in the algorithm uses the representation of the image area given by a modified histogram-of-oriented-gradients descriptor and a support vector machine (SVM) to classify whether the driver, passenger, or no one is interacting with the region of interest. The second step extracts gesture characteristics from temporal dynamics of the features derived in the initial step, which are then inputted to a SVM in order to perform gesture classification from a set of six classes of hand gestures. The rate of correct user classification into one of the three classes is 97.9% on average. Average hand gesture classification rates for the driver and passenger using color and depth input are above 94%. These rates were achieved on in-vehicle collected data over varying illumination conditions and human subjects. This approach demonstrates the feasibility of the hand gesture-based in-vehicle visual user interface. Eshed Ohn-Bar, Cuong Tran 0001, Mohan M. Trivedi |
AutomotiveUI | 1 |