EDBT 2026 Demo / reviewers in the wild / expert
Guy Rosman
dblp:53/3441
· DBLP profile ↗
51ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-9334-1706ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 7 first-author · 21 since 2021Systems, architecture and hardware · 27 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-authorHuman-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | See Something, Say Something: Layered Driver Situational Awareness from Language
Pranay Gupta, Simon Stent, John Gideon, Kayli Battel, Laporsha Dees, Patricio Reyes Gomez, Megan Applegate-Kenton, Emily S. Sumner, Guy Rosman |
IV | 9 |
| 2025 | From Dashboards to Dialogue: Evaluating a Conversational AI Coach for Performance Driving Skill DevelopmentabstractHow can I improve my performance here?Figure 1: After completing a lap, the driver pauses the video at a specific segment and asks ApexTrainer for feedback.The system generates context-aware suggestions based on the referenced location. Jean Marcel dos Reis Costa, Allison Morgan, Hiroshi Yasuda, Emily S. Sumner, Deepak Edakkattil Gopinath, Sheryl Chau, Andrew Best, Guy Rosman, Tiffany L. Chen |
AutomotiveUI | 9 |
| 2025 | Shared Autonomy for Proximal TeachingabstractMotor skills education often requires experienced professionals who can provide personalized instruction. Unfortu-nately, the availability of high-quality training can be limited for specialized tasks, such as high performance racing. Several recent works have proposed AI -assistance for motor skills instruction, ranging from rehabilitation to surgical robot tele-operation. However, these works often make simplifying assumptions on the student learning process, and fail to model how a teacher's assistance interacts with different individuals' abilities when determining optimal teaching strategies. Inspired by the idea of scaffolding from educational psychology, we leverage shared autonomy, a framework for combining user inputs with robot autonomy, to aid with curriculum design. Our key insight is that the way a student's behavior improves in the presence of assistance from an autonomous agent can highlight which sub-skills might be most “learnable” for the student, or within their Zone of Proximal Development. We use this to design Z-COACH, a method for using shared autonomy to provide personalized instruction targeting interpretable task sub-skills. In a user study$(\mathrm{n}=50)$, where we teach high performance racing in a simulated environment of the Thunderhill Raceway Park with the CARLA Autonomous Driving simulator, we show that Z-COACH helps identify which skills each student should first practice, leading to an overall improvement in driving time, behavior, and smoothness. Our work shows that increasingly available semi-autonomous capabilities (e.g. in vehicles, robots) can not only assist human users, but also help teach them. Megha Srivastava, Reihaneh Iranmanesh, Yuchen Cui, Deepak Edakkattil Gopinath, Emily S. Sumner, Andrew Silva, Laporsha Dees, Guy Rosman, Dorsa Sadigh |
HRI | 8 |
| 2025 | ReGen: Generative Robot Simulation via Inverse DesignabstractSimulation plays a key role in scaling robot learning and validating policies, but constructing simulations remains labor-intensive. In this paper, we introduce ReGen, a generative simulation framework that automates this process using inverse design. Given an agent's behavior (such as a motion trajectory or objective function) and its textual description, we infer the underlying scenarios and environments that could have caused the behavior.
Our approach leverages large language models to construct and expand a graph that captures cause-and-effect relationships and relevant entities with properties in the environment, which is then processed to configure a robot simulation environment. Our approach supports (i) augmenting simulations based on ego-agent behaviors, (ii) controllable, counterfactual scenario generation, (iii) reasoning about agent cognition and mental states, and (iv) reasoning with distinct sensing modalities, such as braking due to faulty GPS signals.
We demonstrate our method in autonomous driving and robot manipulation tasks, generating more diverse, complex simulated environments compared to existing simulations with high success rates, and enabling controllable generation for corner cases. This approach enhances the validation of robot policies and supports data or simulation augmentation, advancing scalable robot learning for improved generalization and robustness. Phat Nguyen, Tsun-Hsuan Wang, Zhang-Wei Hong, Erfan Aasi, Andrew Silva, Guy Rosman, Sertac Karaman, Daniela Rus |
ICLR | 6 |
| 2025 | Generating Out-of-Distribution Scenarios Using Language ModelsabstractThe deployment of autonomous vehicles controlled by machine learning techniques requires extensive testing in diverse real-world environments, robust handling of edge cases and out-of-distribution scenarios, and comprehensive safety validation to ensure that these systems can navigate safely and effectively under unpredictable conditions. Addressing Out-OfDistribution (OOD) driving scenarios is essential for enhancing safety, as OOD scenarios help validate the reliability of the models within the vehicle's autonomy stack. However, generating OOD scenarios is challenging due to their long-tailed distribution and rarity in urban driving datasets. Recently, Large Language Models (LLMs) have shown promise in autonomous driving, particularly for their zero-shot generalization and common-sense reasoning capabilities. In this paper, we leverage these LLM strengths to introduce a framework for generating diverse OOD driving scenarios. Our approach uses LLMs to construct a branching tree, where each branch represents a unique OOD scenario. These scenarios are then simulated in the CARLA simulator using an automated framework that aligns scene augmentation with the corresponding textual descriptions. We evaluate our framework through extensive simulations, and assess its performance via a diversity metric that measures the richness of the scenarios. Additionally, we introduce a new “OOD-ness” metric, which quantifies how much the generated scenarios deviate from typical urban driving conditions. Furthermore, we explore the capacity of modern Vision-Language Models (VLMs) to interpret and safely navigate through the simulated OOD scenarios. Our findings offer valuable insights into the reliability of language models in addressing OOD scenarios within the context of urban driving. Erfan Aasi, Phat Nguyen, Shiva Sreeram, Guy Rosman, Sertac Karaman, Daniela Rus |
ICRA | 4 |
| 2025 | Computational Teaching for Driving via Multi-Task Imitation LearningabstractLearning motor skills for sports or performance driving is often done with professional instruction from expert human teachers, whose availability is limited. Our goal is to enable automated teaching via a learned model that interacts with the student similar to a human teacher. However, training such automated teaching systems is limited by the availability of highquality annotated datasets of expert teacher and student interactions as they are difficult to collect at scale. To address this data scarcity problem, we propose an approach for training a coaching system for complex motor tasks such as high performance driving via a Multi-Task Imitation Learning (MTIL) paradigm. MTIL allows our model to learn robust representations by utilizing self-supervised training signals from more readily available non-interactive datasets of humans performing the task of interest. We validate our approach with (1) a semi-synthetic dataset created from real human driving trajectories, (2) a professional track driving instruction dataset, (3) a track-racing driving simulator human-subject study, and (4) a system demonstration on an instrumented car at a race track. Our experiments show that the right set of auxiliary machine learning tasks improves prediction of teaching instructions. Moreover, in the human subjects study, students exposed to the instructions from our teaching system improve their ability to stay within track limits, and show favorable perception of the model's interaction with them, in terms of usefulness and satisfaction. Deepak Edakkattil Gopinath, Xiongyi Cui, Jonathan A. DeCastro, Emily S. Sumner, Jean Costa, Hiroshi Yasuda, Allison Morgan, Laporsha Dees, Sheryl Chau, John J. Leonard, Tiffany L. Chen, Guy Rosman, Avinash Balachandran |
ICRA | 12 |
| 2025 | Think Deep and Fast: Learning Neural Nonlinear Opinion Dynamics from Inverse Dynamic Games for Split-Second InteractionsabstractNon-cooperative interactions commonly occur in multi-agent scenarios such as car racing, where an ego vehicle can choose to overtake the rival, or stay behind it until a safe overtaking “corridor” opens. While an expert human can do well at making such time-sensitive decisions, autonomous agents are incapable of rapidly reasoning about complex, potentially conflicting options, leading to suboptimal behaviors such as deadlocks. Recently, the nonlinear opinion dynamics (NOD) model has proven to exhibit fast opinion formation and avoidance of decision deadlocks. However, NOD modeling parameters are oftentimes assumed fixed, limiting their applicability in complex and dynamic environments. It remains an open challenge to determine such parameters automatically and adaptively, accounting for the ever-changing environment. In this work, we propose for the first time a learning-based and game-theoretic approach to synthesize a Neural NOD model from expert demonstrations, given as a dataset containing (possibly incomplete) state and action trajectories of interacting agents. We demonstrate Neural NOD's ability to make fast and deadlock-free decisions in a simulated autonomous racing example. We find that Neural NOD consistently outperforms the state-of-the-art data-driven inverse game baseline in terms of safety and overtaking performance. Haimin Hu, Jaime Fernández Fisac, Naomi Ehrich Leonard, Deepak Edakkattil Gopinath, Jonathan A. DeCastro, Guy Rosman |
ICRA | 6 |
| 2025 | Hypergraph-Transformer (HGT) for Interaction Event Prediction in Laparoscopic and Robotic SurgeryabstractUnderstanding and anticipating events and actions is critical for intraoperative assistance and decision-making during minimally invasive surgery. We propose a predictive neural network that is capable of understanding and predicting critical interaction aspects of surgical workflow based on endoscopic, intracorporeal video data, while flexibly leveraging surgical knowledge graphs. The approach incorporates a hypergraph-transformer (HGT) structure that encodes expert knowledge into the network design and predicts the hidden embedding of the graph. We verify our approach on established surgical datasets and applications, including the prediction of action-triplets, and the achievement of the Critical View of Safety (CVS), which is a critical safety measure. Moreover, we address specific, safety-related forecasts of surgical processes, such as predicting the clipping of the cystic duct or artery without prior achievement of the CVS. Our results demonstrate improvement in prediction of interactive event when incorporating with our approach compared to unstructured alternatives. Lianhao Yin, Yutong Ban, Jennifer A. Eckhoff, Ozanan R. Meireles, Daniela Rus, Guy Rosman |
ICRA | 6 |
| 2025 | Beyond Breathalyzers: Towards Pre-Driving Sobriety Testing with a Driver Monitoring CameraabstractField sobriety tests and breathalyzers are commonly used to prevent alcohol-impaired driving, but are expensive and time-consuming to administer. We propose a set of sobriety tests which, in contrast, can feasibly be automated and deployed to modern vehicles equipped with a driver monitoring camera. Our tests are inspired by research on the physiological effects of alcohol, with particular focus on eye movements and gaze behavior. We run an exploratory in-lab study with N=50 subjects (20 alcohol-impaired, 30 control), and train a variety of models to detect alcohol impairment. We find that, using only 10 seconds of observations of the driver, one of the four proposed tests performs comparably to existing non-breathalyzer field sobriety tests. We make our code and data available to support further research efforts to combat alcohol-impaired driving: https://toyotaresearchinstitute.github.io/IV25-beyond-breathalysers/. Simon Stent, John Gideon, Kimimasa Tamura, Avinash Balachandran, Guy Rosman |
IV | 5 |
| 2025 | Cognitive Distraction Detection Using Gaze and Pupil with an Interpretable ApproachabstractCognitive distraction (CD) is one of the major causes of traffic accidents, but there remains room to improve its detection. Most prior research on CD detection has commonly used basic statistical measures (e.g., mean, standard deviation) of driver-facing camera signals such as gaze and pupil size. However, these signals often exhibit subtle and complex patterns that conventional approaches cannot fully capture. In this paper, we evaluate a wide range of machine learning models and feature extraction methods using data from 52 participants in a driving simulator under two cognitive distraction inducing tasks (n-back and statement tasks). Our results demonstrate that combining gaze, pupil, and features derived from physiological signals (e.g., fixation saccade ratio and gaze entropy) and comprehensive time-series feature extraction boosts detection performance. While deep neural networks (Transformers) excel at modeling intricate relationships, our results show that tree-based ensemble methods (e.g., CatBoost) achieve comparable or higher detection performance while maintaining their advantage of better interpretability. Cross-task experiments further show that models trained on one type of task can generalize to another task. Feature analyses (via SHAP and Sobol) reveal that nonlinearity in vertical gaze movements, baseline pupil size, and greater minimum gaze distance are related to CD. These findings suggest that integrating multiple modalities, sophisticated feature engineering, and employing models capable of capturing nonlinear interactions are effective strategies for detecting CD. To support future research in this field, we release our code and preprocessed data: https://toyotaresearchinstitute.github.io/IV25-cognitive-distraction/. Kimimasa Tamura, Simon Stent, John Gideon, Kohei Shintani, Guy Rosman |
IV | 5 |
| 2025 | Estimating cognitive biases with attention-aware inverse planningabstractPeople's goal-directed behaviors are influenced by their cognitive biases, and autonomous systems that interact with people should be aware of this. For example, people's attention to objects in their environment will be biased in a way that systematically affects how they perform everyday tasks such as driving to work. Here, building on recent work in computational cognitive science, we formally articulate the \textit{attention-aware inverse planning problem}, in which the goal is to estimate a person's attentional biases from their actions. We demonstrate how attention-aware inverse planning systematically differs from standard inverse reinforcement learning and how cognitive biases can be inferred from behavior. Finally, we present an approach to attention-aware inverse planning that combines deep reinforcement learning with computational cognitive modeling. We use this approach to infer the attentional strategies of RL agents in real-life driving scenarios selected from the Waymo Open Dataset, demonstrating the scalability of estimating cognitive biases with attention-aware inverse planning. Sounak Banerjee 0002, Daphne Cornelisse, Deepak Edakkattil Gopinath, Emily S. Sumner, Jonathan A. DeCastro, Guy Rosman, Eugene Vinitsky, Mark K. Ho |
NeurIPS | 6 |
| 2024 | Drive Anywhere: Generalizable End-to-end Autonomous Driving with Multi-modal Foundation ModelsabstractAs autonomous driving technology matures, end-to-end methodologies have emerged as a leading strategy, promising seamless integration from perception to control via deep learning. However, existing systems grapple with challenges such as unexpected open set environments and the complexity of black-box models. At the same time, the evolution of deep learning introduces larger, multimodal foundational models, offering multi-modal visual and textual understanding. In this paper, we harness these multimodal foundation models to enhance the robustness and adaptability of autonomous driving systems. We introduce a method to extract nuanced spatial features from transformers and the incorporation of latent space simulation for improved training and policy debugging. We use pixel/patch-aligned feature descriptors to expand foundational model capabilities to create an end-to-end multimodal driving model, demonstrating unparalleled results in diverse tests. Our solution combines language with visual perception and achieves significantly greater robustness on out-of-distribution situations. Tsun-Hsuan Wang, Alaa Maalouf, Wei Xiao 0003, Yutong Ban, Alexander Amini, Guy Rosman, Sertac Karaman, Daniela Rus |
ICRA | 6 |
| 2024 | Learning autonomous driving from aerial imageryabstractIn this work, we consider the problem of learning end to end perception to control for ground vehicles solely from aerial imagery. Photogrammetric simulators allow the synthesis of novel views through the transformation of pre-generated assets into novel views. However, they have a large setup cost, require careful collection of data and often human effort to create usable simulators. We use a Neural Radiance Field (NeRF) as an intermediate representation to synthesize novel views from the point of view of a ground vehicle. These novel viewpoints can then be used for several downstream autonomous navigation applications. In this work, we demonstrate the utility of novel view synthesis though the application of training a policy for end to end learning from images and depth data. In a traditional real to sim to real framework, the collected data would be transformed into a visual simulator which could then be used to generate novel views. In contrast, using a NeRF allows a compact representation and the ability to optimize over the parameters of the visual simulator as more data is gathered in the environment. We demonstrate the efficacy of our method in a custom built mini-city environment through the deployment of imitation policies on robotic cars. We additionally consider the task of place localization and demonstrate that our method is able to relocalize the car in the real world. Varun Murali, Guy Rosman, Sertac Karaman, Daniela Rus |
IROS | 2 |
| 2024 | Online Adaptation of Learned Vehicle Dynamics Model with Meta-Learning ApproachabstractWe represent a vehicle dynamics model for autonomous driving near the limits of handling via a multilayer neural network. Online adaptation is desirable in order to address unseen environments. However, the model needs to adapt to new environments without forgetting previously encountered ones. In this study, we apply Continual-MAML to overcome this difficulty. It enables the model to adapt to the previously encountered environments quickly and efficiently by starting updates from optimized initial parameters. We evaluate the impact of online model adaptation with respect to inference performance and impact on control performance of a model predictive path integral (MPPI) controller using the TRIKart platform. The neural network was pre-trained using driving data collected in our test environment, and experiments for online adaptation were executed on multiple different road conditions not contained in the training data. Empirical results show that the model using Continual-MAML outperforms the fixed model and the model using gradient descent in test set loss and online tracking performance of MPPI. Yuki Tsuchiya, Thomas Balch, Paul Drews, Guy Rosman |
IROS | 4 |
| 2024 | Can Pupillometry be used to Detect Driver Hazard Awareness?abstractModern Advanced Driver-Assistance Systems (ADAS) increasingly rely on interactions between vehicle and human driver. To inform these interactions, it is helpful for a vehicle system to have a good understanding of a driver's situational awareness. In this work we explore a relatively under-exploited, passively measurable signal which might provide insight into a driver's awareness: the constriction and dilation of their pupils over time, or pupillometry. We ask whether pupillometry might be practically useful to detect if and when a driver becomes aware of a road hazard. Using a dataset of driver responses to both hazardous and routine scenarios during simulated semi-automated driving, we compare models trained on pupillometric data to a model trained on facial responses, and demonstrate how their performances differ in terms of accuracy and latency. While a driver's facial expressions are, as expected, a useful cue to determine awareness (0.82 AUC on held-out test stimuli), we find that pupillometric data alone can provide an even more meaningful signal (0.93 AUC). In addition, we find that the pupillometric model performance degrades more gracefully than the face model when tested on unseen subjects, while fusing models yields further accuracy and latency improvements given sufficient training data. We characterize the shape of the performance vs. latency curve for all models and make our code available for reproducibility. Kimimasa Tamura, John Gideon, Simon Stent, Guy Rosman |
SMC | 4 |
| 2024 | Concept Graph Neural Networks for Surgical Video UnderstandingabstractAnalysis of relations between objects and comprehension of abstract concepts in the surgical video is important in AI-augmented surgery. However, building models that integrate our knowledge and understanding of surgery remains a challenging endeavor. In this paper, we propose a novel way to integrate conceptual knowledge into temporal analysis tasks using temporal concept graph networks. In the proposed networks, a knowledge graph is incorporated into the temporal video analysis of surgical notions, learning the meaning of concepts and relations as they apply to the data. We demonstrate results in surgical video data for tasks such as verification of the critical view of safety, estimation of the Parkland grading scale as well as recognizing instrument-action-tissue triplets. The results show that our method improves the recognition and detection of complex benchmarks as well as enables other analytic applications of interest. Yutong Ban, Jennifer A. Eckhoff, Thomas M. Ward, Daniel A. Hashimoto, Ozanan R. Meireles, Daniela Rus, Guy Rosman |
IEEE Trans. Medical Imaging | 7 |
| 2023 | MPOGames: Efficient Multimodal Partially Observable Dynamic GamesabstractGame theoretic methods have become popular for planning and prediction in situations involving rich multi-agent interactions. However, these methods often assume the existence of a single local Nash equilibria and are hence unable to handle uncertainty in the intentions of different agents. While maximum entropy (MaxEnt) dynamic games try to address this issue, practical approaches solve for MaxEnt Nash equilibria using linear-quadratic approximations which are restricted to unimodal responses and unsuitable for scenarios with multiple local Nash equilibria. By reformulating the problem as a POMDP, we propose MPOGames, a method for efficiently solving MaxEnt dynamic games that captures the interactions between local Nash equilibria. We show the importance of uncertainty-aware game theoretic methods via a two-agent merge case study. Finally, we prove the real-time capabilities of our approach with hardware experiments on a 1/10th scale car platform. Oswin So, Paul Drews, Thomas Balch, Velin D. Dimitrov, Guy Rosman, Evangelos A. Theodorou |
ICRA | 5 |
| 2022 | A Deep Concept Graph Network for Interaction-Aware Trajectory PredictionabstractTemporal patterns (how vehicles behave in our observed past) underline our reasoning of how people drive on the road, and can explain why we make certain predictions about interactions among road agents. In this paper we propose the ConceptNet trajectory predictor - a novel prediction framework that is able to incorporate agent interactions as explicit edges in a temporal knowledge graph. We demonstrate the sample efficiency and the overall accuracy of the proposed approach, and show that using the graphical structure to explicitly model interactions enables better detection of agent interactions and improved trajectory predictions on a large real-world driving dataset. Yutong Ban, Xiao Li 0025, Guy Rosman, Igor Gilitschenski, Ozanan R. Meireles, Sertac Karaman, Daniela Rus |
ICRA | 3 |
| 2022 | Leveraging Smooth Attention Prior for Multi-Agent Trajectory PredictionabstractMulti-agent interactions are important to model for forecasting other agents' behaviors and trajectories. At a certain time, to forecast a reasonable future trajectory, each agent needs to pay attention to the interactions with only a small group of most relevant agents instead of unnecessarily paying attention to all the other agents. However, existing attention modeling works ignore that human attention in driving does not change rapidly, and may introduce fluctuating attention across time steps. In this paper, we formulate an attention model for multi-agent interactions based on a total variation temporal smoothness prior and propose a trajectory prediction architecture that leverages the knowledge of these attended interactions. We demonstrate how the total variation attention prior along with the new sequence prediction loss terms leads to smoother attention and more sample-efficient learning of multi-agent trajectory prediction, and show its advantages in terms of prediction accuracy by comparing it with the state-of-the-art approaches on both synthetic and naturalistic driving data. We demonstrate the performance of our algorithm for trajectory prediction on the INTERACTION dataset on our website11https://sites.google.com/view/smoothness-attention. Zhangjie Cao, Erdem Biyik, Guy Rosman, Dorsa Sadigh |
ICRA | 3 |
| 2022 | HYPER: Learned Hybrid Trajectory Prediction via Factored Inference and Adaptive SamplingabstractModeling multi-modal high-level intent is important for ensuring diversity in trajectory prediction. Existing approaches explore the discrete nature of human intent before predicting continuous trajectories, to improve accuracy and support explainability. However, these approaches often assume the intent to remain fixed over the prediction horizon, which is problematic in practice, especially over longer horizons. To overcome this limitation, we introduce HYPER, a general and expressive hybrid prediction framework that models evolving human intent. By modeling traffic agents as a hybrid discrete-continuous system, our approach is capable of predicting discrete intent changes over time. We learn the probabilistic hybrid model via a maximum likelihood estimation problem and leverage neural proposal distributions to sample adaptively from the exponentially growing discrete space. The overall approach affords a better trade-off between accuracy and coverage. We train and validate our model on the Argoverse dataset, and demonstrate its effectiveness through comprehensive ablation studies and comparisons with state-of-the-art models. Xin Huang 0018, Guy Rosman, Igor Gilitschenski, Ashkan Jasour, Stephen G. McGill, John J. Leonard, Brian C. Williams |
ICRA | 2 |
| 2022 | Trajectory Prediction with Linguistic RepresentationsabstractLanguage allows humans to build mental models that interpret what is happening around them resulting in more accurate long-term predictions. We present a novel trajectory prediction model that uses linguistic intermediate representations to forecast trajectories, and is trained using trajectory samples with partially-annotated captions. The model learns the meaning of each of the words without direct per-word supervision. At inference time, it generates a linguistic description of trajectories which captures maneuvers and interactions over an extended time interval. This generated description is used to refine predictions of the trajectories of multiple agents. We train and validate our model on the Argoverse dataset, and demonstrate improved accuracy results in trajectory prediction. In addition, our model is more interpretable: it presents part of its reasoning in plain language as captions, which can aid model development and can aid in building confidence in the model before deploying it. Yen-Ling Kuo, Xin Huang 0018, Andrei Barbu, Stephen G. McGill, Boris Katz, John J. Leonard, Guy Rosman |
ICRA | 7 |
| 2022 | TIP: Task-Informed Motion Prediction for Intelligent VehiclesabstractWhen predicting trajectories of road agents, motion predictors often approximate the future distribution by a limited number of samples. This constraint requires the predictors to generate samples that best support the task given task specifications. However, existing predictors are often optimized and evaluated via task-agnostic measures without accounting for the use of predictions in downstream tasks, and thus could result in sub-optimal task performance. In this paper, we propose a task-informed motion prediction model that better supports the tasks through its predictions by jointly reasoning about prediction accuracy and the utility of the downstream tasks during training. The task utility function is commonly used to evaluate task performance. It does not require the full task information, but rather a specification of the utility of the task, resulting in predictors that are tailored to different downstream tasks. We demonstrate our approach on two use cases of common decision making tasks and their utility functions, in the context of autonomous driving and parallel autonomy. Experiment results show that our predictor produces accurate predictions that improve the task performance by a large margin in both tasks when compared to task-agnostic baselines on the Waymo Open Motion dataset. Xin Huang 0018, Guy Rosman, Ashkan Jasour, Stephen G. McGill, John J. Leonard, Brian C. Williams |
IROS | 2 |
| 2021 | Aggregating Long-Term Context for Learning Laparoscopic and Robot-Assisted Surgical WorkflowsabstractAnalyzing surgical workflow is crucial for surgical assistance robots to understand surgeries. With the understanding of the complete surgical workflow, the robots are able to assist the surgeons in intra-operative events, such as by giving a warning when the surgeon is entering specific keys or high-risk phases. Deep learning techniques have recently been widely applied to recognizing surgical workflows. Many of the existing temporal neural network models are limited in their capability to handle long-term dependencies in the data, instead, relying upon the strong performance of the underlying per-frame visual models. We propose a new temporal network structure that leverages task-specific network representation to collect long-term sufficient statistics that are propagated by a sufficient statistics model (SSM). We implement our approach within an LSTM backbone for the task of surgical phase recognition and explore several choices for propagated statistics. We demonstrate superior results over existing and novel state-of-the-art segmentation techniques on two laparoscopic cholecystectomy datasets: the publicly available Cholec80 dataset and MGH100, a novel dataset with more challenging and clinically meaningful segment labels. Yutong Ban, Guy Rosman, Thomas M. Ward, Daniel A. Hashimoto, Taisei Kondo, Hidekazu Iwaki, Ozanan R. Meireles, Daniela Rus |
ICRA | 2 |
| 2021 | Risk Conditioned Neural Motion PlanningabstractRisk-bounded motion planning is an important yet difficult problem for safety-critical tasks. While existing mathematical programming methods offer theoretical guarantees in the context of constrained Markov decision processes, they either lack scalability in solving larger problems or produce conservative plans. Recent advances in deep reinforcement learning improve scalability by learning policy networks as function approximators. In this paper, we propose an extension of soft actor critic model to estimate the execution risk of a plan through a risk critic and produce risk-bounded policies efficiently by adding an extra risk term in the loss function of the policy network. We define the execution risk in an accurate form, as opposed to approximating it through a summation of immediate risks at each time step that leads to conservative plans. Our proposed model is conditioned on a continuous spectrum of risk bounds, allowing the user to adjust the risk-averse level of the agent on the fly. Through a set of experiments, we show the advantage of our model in terms of both computational time and plan quality, compared to a state-of-the-art mathematical programming baseline, and validate its performance in more complicated scenarios, including nonlinear dynamics and larger state space. Xin Huang 0018, Meng Feng, Ashkan Jasour, Guy Rosman, Brian C. Williams |
IROS | 4 |
| 2020 | Driving Through Ghosts: Behavioral Cloning with False PositivesabstractSafe autonomous driving requires robust detection of other traffic participants. However, robust does not mean perfect, and safe systems typically minimize missed detections at the expense of a higher false positive rate. This results in conservative and yet potentially dangerous behavior such as avoiding imaginary obstacles. In the context of behavioral cloning, perceptual errors at training time can lead to learning difficulties or wrong policies, as expert demonstrations might be inconsistent with the perceived world state. In this work, we propose a behavioral cloning approach that can safely leverage imperfect perception without being conservative. Our core contribution is a novel representation of perceptual uncertainty for learning to plan. We propose a new probabilistic birds-eye-view semantic grid to encode the noisy output of object perception systems. We then leverage expert demonstrations to learn an imitative driving policy using this probabilistic representation. Using the CARLA simulator, we show that our approach can safely overcome critical false positives that would otherwise lead to catastrophic failures or conservative behavior. Andreas Bühler, Adrien Gaidon, Andrei Cramariuc, Rares Ambrus, Guy Rosman, Wolfram Burgard |
IROS | 5 |
| 2020 | Behaviorally Diverse Traffic Simulation via Reinforcement LearningabstractTraffic simulators are important tools in autonomous driving development. While continuous progress has been made to provide developers more options for modeling various traffic participants, tuning these models to increase their behavioral diversity while maintaining quality is often very challenging. This paper introduces an easily-tunable policy generation algorithm for autonomous driving agents. The proposed algorithm balances diversity and driving skills by leveraging the representation and exploration abilities of deep reinforcement learning via a distinct policy set selector. Moreover, we present an algorithm utilizing intrinsic rewards to widen behavioral differences in the training. To provide quantitative assessments, we develop two trajectory-based evaluation metrics which measure the differences among policies and behavioral coverage. We experimentally show the effectiveness of our methods on several challenging intersection scenes. Shinya Shiroshita, Shirou Maruyama, Daisuke Nishiyama, Mario Ynocente Castro, Karim Hamzaoui, Guy Rosman, Jonathan A. DeCastro, Kuan-Hui Lee, Adrien Gaidon |
IROS | 6 |
| 2019 | Variational End-to-End Navigation and LocalizationabstractDeep learning has revolutionized the ability to learn “end-to-end” autonomous vehicle control directly from raw sensory data. While there have been recent extensions to handle forms of navigation instruction, these works are unable to capture the full distribution of possible actions that could be taken and to reason about localization of the robot within the environment. In this paper, we extend end-to-end driving networks with the ability to perform point-to-point navigation as well as probabilistic localization using only noisy GPS data. We define a novel variational network capable of learning from raw camera data of the environment as well as higher level roadmaps to predict (1) a full probability distribution over the possible control commands; and (2) a deterministic control command capable of navigating on the route specified within the map. Additionally, we formulate how our model can be used to localize the robot according to correspondences between the map and the observed visual road topology, inspired by the rough localization that human drivers can perform. We test our algorithms on real-world driving data that the vehicle has never driven through before, and integrate our point-topoint navigation algorithms onboard a full-scale autonomous vehicle for real-time performance. Our localization algorithm is also evaluated over a new set of roads and intersections to demonstrates rough pose localization even in situations without any GPS prior. Alexander Amini, Guy Rosman, Sertac Karaman, Daniela Rus |
ICRA | 2 |
| 2019 | Uncertainty-Aware Driver Trajectory Prediction at Urban IntersectionsabstractPredicting the motion of a driver’s vehicle is crucial for advanced driving systems, enabling detection of potential risks towards shared control between the driver and automation systems. In this paper, we propose a variational neural network approach that predicts future driver trajectory distributions for the vehicle based on multiple sensors.Our predictor generates both a conditional variational distribution of future trajectories, as well as a confidence estimate for different time horizons. Our approach allows us to handle inherently uncertain situations, and reason about information gain from each input, as well as combine our model with additional predictors, creating a mixture of experts.We show how to augment the variational predictor with a physics-based predictor, and based on their confidence estimations, improve overall system performance. The resulting combined model is aware of the uncertainty associated with its predictions, which can help the vehicle autonomy to make decisions with more confidence. The model is validated on real-world urban driving data collected in multiple locations. This validation demonstrates that our approach improves the prediction error of a physics-based model by 25% while successfully identifying the uncertain cases with 82% accuracy. Xin Huang 0018, Stephen G. McGill, Brian C. Williams, Luke Fletcher, Guy Rosman |
ICRA | 5 |
| 2019 | Infrastructure-free NLoS Obstacle Detection for Autonomous CarsabstractCurrent perception systems mostly require direct line of sight to anticipate and ultimately prevent potential collisions at intersections with other road users. We present a fully integrated autonomous system capable of detecting shadows or weak illumination changes on the ground caused by a dynamic obstacle in NLoS scenarios. This additional virtual sensor “ShadowCam” extends the signal range utilized so far by computer-vision ADASs. We show that (1) our algorithm maintains the mean classification accuracy of around 70% even when it doesn't rely on infrastructure - such as AprilTags - as an image registration method. We validate (2) in real-world experiments that our autonomous car driving in night time conditions detects a hidden approaching car earlier with our virtual sensor than with the front facing 2-D LiDAR. Felix Naser, Igor Gilitschenski, Alexander Amini, Christina Liao, Guy Rosman, Sertac Karaman, Daniela Rus |
IROS | 5 |
| 2018 | Task-Specific Sensor Planning for Robotic Assembly TasksabstractWhen performing multi-robot tasks, sensory feedback is crucial in reducing uncertainty for correct execution. Yet the utilization of sensors should be planned as an integral part of the task planning, taken into account several factors such as the tolerance of different inferred properties of the scene and interaction with different agents. In this paper we handle this complex problem in a principled, yet efficient way. We use surrogate predictors based on open-loop simulation to estimate and bound the probability of success for specific tasks. We reason about such task-specific uncertainty approximants and their effectiveness. We show how they can be incorporated into a multi-robot planner, and demonstrate results with a team of robots performing assembly tasks. Guy Rosman, Changhyun Choi, Mehmet Remzi Dogar, John W. Fisher III, Daniela Rus |
ICRA | 1 |
| 2018 | Variational Autoencoder for End-to-End Control of Autonomous Driving with Novelty Detection and Training De-biasingabstractThis paper introduces a new method for end-to-end training of deep neural networks (DNNs) and evaluates it in the context of autonomous driving. DNN training has been shown to result in high accuracy for perception to action learning given sufficient training data. However, the trained models may fail without warning in situations with insufficient or biased training data. In this paper, we propose and evaluate a novel architecture for self-supervised learning of latent variables to detect the insufficiently trained situations. Our method also addresses training data imbalance, by learning a set of underlying latent variables that characterize the training data and evaluate potential biases. We show how these latent distributions can be leveraged to adapt and accelerate the training pipeline by training on only a fraction of the total dataset. We evaluate our approach on a challenging dataset for driving. The data is collected from a full-scale autonomous vehicle. Our method provides qualitative explanation for the latent variables learned in the model. Finally, we show how our model can be additionally trained as an end-to-end controller, directly outputting a steering control command for an autonomous vehicle. Alexander Amini, Wilko Schwarting, Guy Rosman, Brandon Araki, Sertac Karaman, Daniela Rus |
IROS | 3 |
| 2018 | The Manhattan Frame Model - Manhattan World Inference in the Space of Surface NormalsabstractObjects and structures within man-made environments typically exhibit a high degree of organization in the form of orthogonal and parallel planes. Traditional approaches utilize these regularities via the restrictive, and rather local, Manhattan World (MW) assumption which posits that every plane is perpendicular to one of the axes of a single coordinate system. The aforementioned regularities are especially evident in the surface normal distribution of a scene where they manifest as orthogonally-coupled clusters. This motivates the introduction of the Manhattan-Frame (MF) model which captures the notion of an MW in the surface normals space, the unit sphere, and two probabilistic MF models over this space. First, for a single MF we propose novel real-time MAP inference algorithms, evaluate their performance and their use in drift-free rotation estimation. Second, to capture the complexity of real-world scenes at a global scale, we extend the MF model to a probabilistic mixture of Manhattan Frames (MMF). For MMF inference we propose a simple MAP inference algorithm and an adaptive Markov-Chain Monte-Carlo sampling algorithm with Metropolis-Hastings split/merge moves that let us infer the unknown number of mixture components. We demonstrate the versatility of the MMF model and inference algorithm across several scales of man-made environments. Julian Straub, Oren Freifeld, Guy Rosman, John J. Leonard, John W. Fisher III |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Persistent surveillance of events with unknown, time-varying statisticsabstractWe consider the problem of monitoring stochastic, time-varying events occurring at discrete locations. Our problem formulation extends prior work in persistent surveillance by considering the objective of maximizing event detections in unknown, dynamic environments where the rates of events are time-inhomogeneous and may be subject to abrupt changes. We propose a novel monitoring algorithm that effectively strikes a balance between exploration and exploitation as well as a balance between remembering and discarding information to handle temporal variations in unknown environments. We present an analysis proving the long-run average optimality of the policies generated by our algorithm under the assumption that the total temporal variations are sub-linear. We present simulation results demonstrating the effectiveness of our algorithm in several monitoring scenarios inspired by real-world applications, and its robustness to both continuous-random and abrupt changes in the statistics of the observed processes. Cenk Baykal, Guy Rosman, Sebastian Claici, Daniela Rus |
ICRA | 2 |
| 2017 | Duckietown: An open, inexpensive and flexible platform for autonomy education and researchabstractDuckietown is an open, inexpensive and flexible platform for autonomy education and research. The platform comprises small autonomous vehicles (“Duckiebots”) built from off-the-shelf components, and cities (“Duckietowns”) complete with roads, signage, traffic lights, obstacles, and citizens (duckies) in need of transportation. The Duckietown platform offers a wide range of functionalities at a low cost. Duckiebots sense the world with only one monocular camera and perform all processing onboard with a Raspberry Pi 2, yet are able to: follow lanes while avoiding obstacles, pedestrians (duckies) and other Duckiebots, localize within a global map, navigate a city, and coordinate with other Duckiebots to avoid collisions. Duckietown is a useful tool since educators and researchers can save money and time by not having to develop all of the necessary supporting infrastructure and capabilities. All materials are available as open source, and the hope is that others in the community will adopt the platform for education and research. Liam Paull, Jacopo Tani, Heejin Ahn, Javier Alonso-Mora, Luca Carlone, Michal Cáp, Yu Fan Chen, Changhyun Choi, Jeff Dusek, Yajun Fang, Daniel Hoehener, Shih-Yuan Liu, Michael Novitzky, Igor Franzoni Okuyama, Jason Pazis, Guy Rosman, Valerio Varricchio, Hsueh-Cheng Wang, Dmitry S. Yershov, Hang Zhao 0021, Michael Benjamin, Christopher Carr, Maria T. Zuber, Sertac Karaman, Emilio Frazzoli, Domitilla Del Vecchio, Daniela Rus, Jonathan P. How, John J. Leonard, Andrea Censi |
ICRA | 16 |
| 2017 | Machine learning and coresets for automated real-time video segmentation of laparoscopic and robot-assisted surgeryabstractContext-aware segmentation of laparoscopic and robot assisted surgical video has been shown to improve performance and perioperative workflow efficiency, and can be used for education and time-critical consultation. Modern pressures on productivity preclude manual video analysis, and hospital policies and legacy infrastructure are often prohibitive of recording and storing large amounts of data. In this paper we present a system that automatically generates a video segmentation of laparoscopic and robot-assisted procedures according to their underlying surgical phases using minimal computational resources, and low amounts of training data. Our system uses an SVM and HMM in combination with an augmented feature space that captures the variability of these video streams without requiring analysis of the nonrigid and variable environment. By using the data reduction capabilities of online k-segment coreset algorithms we can efficiently produce results of approximately equal quality, in realtime. We evaluate our system in cross-validation experiments and propose a blueprint for piloting such a system in a real operating room environment with minimal risk factors. Mikhail Volkov 0002, Daniel A. Hashimoto, Guy Rosman, Ozanan R. Meireles, Daniela Rus |
ICRA | 3 |
| 2017 | Hybrid control and learning with coresets for autonomous vehiclesabstractModern autonomous systems such as driverless vehicles need to safely operate in a wide range of conditions. A potential solution is to employ a hybrid systems approach, where safety is guaranteed in each individual mode within the system. This offsets complexity and responsibility from the individual controllers onto the complexity of determining discrete mode transitions. In this work we propose an efficient framework based on recursive neural networks and coreset data summarization to learn the transitions between an arbitrary number of controller modes that can have arbitrary complexity. Our approach allows us to efficiently gather annotation data from the large-scale datasets that are required to train such hybrid nonlinear systems to be safe under all operating conditions, favoring underexplored parts of the data. We demonstrate the construction of the embedding, and efficient detection of switching points for autonomous and non-autonomous car data. We further show how our approach enables efficient sampling of training data, to further improve either our embedding or the controllers. Guy Rosman, Liam Paull, Daniela Rus |
IROS | 1 |
| 2016 | Real-Time Depth Refinement for Specular ObjectsabstractThe introduction of consumer RGB-D scanners set off a major boost in 3D computer vision research. Yet, the precision of existing depth scanners is not accurate enough to recover fine details of a scanned object. While modern shading based depth refinement methods have been proven to work well with Lambertian objects, they break down in the presence of specularities. We present a novel shape from shading framework that addresses this issue and enhances both diffuse and specular objects' depth profiles. We take advantage of the built-in monochromatic IR projector and IR images of the RGB-D scanners and present a lighting model that accounts for the specular regions in the input image. Using this model, we reconstruct the depth map in real-time. Both quantitative tests and visual evaluations prove that the proposed method produces state of the art depth reconstruction results. Roy Or-El, Rom Hershkovitz, Aaron Wetzler, Guy Rosman, Alfred M. Bruckstein, Ron Kimmel |
CVPR | 4 |
| 2016 | Information-Driven Adaptive Structured-Light ScannersabstractSensor planning and active sensing, long studied in robotics, adapt sensor parameters to maximize a utility function while constraining resource expenditures. Here we consider information gain as the utility function. While these concepts are often used to reason about 3D sensors, these are usually treated as a predefined, black-box, component. In this paper we show how the same principles can be used as part of the 3D sensor. We describe the relevant generative model for structured-light 3D scanning and show how adaptive pattern selection can maximize information gain in an open-loop-feedback manner. We then demonstrate how different choices of relevant variable sets (corresponding to the subproblems of locatization and mapping) lead to different criteria for pattern selection and can be computed in an online fashion. We show results for both subproblems with several pattern dictionary choices and demonstrate their usefulness for pose estimation and depth acquisition. Guy Rosman, Daniela Rus, John W. Fisher III |
CVPR | 1 |
| 2016 | Persistent Surveillance of Events with Unknown Rate Statistics
Cenk Baykal, Guy Rosman, Kyle Kotowick, Mark Donahue, Daniela Rus |
WAFR | 2 |
| 2015 | RGBD-fusion: Real-time high precision depth recoveryabstractThe popularity of low-cost RGB-D scanners is increasing on a daily basis. Nevertheless, existing scanners often cannot capture subtle details in the environment. We present a novel method to enhance the depth map by fusing the intensity and depth information to create more detailed range profiles. The lighting model we use can handle natural scene illumination. It is integrated in a shape from shading like technique to improve the visual fidelity of the reconstructed object. Unlike previous efforts in this domain, the detailed geometry is calculated directly, without the need to explicitly find and integrate surface normals. In addition, the proposed method operates four orders of magnitude faster than the state of the art. Qualitative and quantitative visual and statistical evidence support the improvement in the depth obtained by the suggested method. Roy Or-El, Guy Rosman, Aaron Wetzler, Ron Kimmel, Alfred M. Bruckstein |
CVPR | 2 |
| 2015 | Coresets for visual summarization with applications to loop closureabstractIn continuously operating robotic systems, efficient representation of the previously seen camera feed is crucial. Using a highly efficient compression coreset method, we formulate a new method for hierarchical retrieval of frames from large video streams collected online by a moving robot. We demonstrate how to utilize the resulting structure for efficient loop-closure by a novel sampling approach that is adaptive to the structure of the video. The same structure also allows us to create a highly-effective search tool for large-scale videos, which we demonstrate in this paper. We show the efficiency of proposed approaches for retrieval and loop closure on standard datasets, and on a large-scale video from a mobile camera. Mikhail Volkov 0002, Guy Rosman, Dan Feldman, John W. Fisher III, Daniela Rus |
ICRA | 2 |
| 2015 | Fleye on the car: big data meets the internet of thingsabstractVehicle-based vision algorithms, such as the collision alert systems [4], are able to interpret a scene in real-time and provide drivers with immediate feedback. However, such technologies are based on cameras on the car, limited to the vicinity of the car, severely limiting their potential. They cannot find empty parking slots, bypass traffic jams, or warn about dangers outside the car's immediate surrounding. An intelligent driving system augmented with additional sensors and network inputs may significantly reduce the number of accidents, improve traffic congestion, and care for the safety and quality of people's lives. Soliman Nasser, Andew Barry, Marek Doniec, Guy Peled, Guy Rosman, Daniela Rus, Mikhail Volkov 0002, Dan Feldman |
IPSN | 5 |
| 2015 | Multi-Region Active Contours with a Single Level Set FunctionabstractSegmenting an image into an arbitrary number of coherent regions is at the core of image understanding. Many formulations of the segmentation problem have been suggested over the past years. These formulations include, among others, axiomatic functionals, which are hard to implement and analyze, and graph-based alternatives, which impose a non-geometric metric on the problem. We propose a novel method for segmenting an image into an arbitrary number of regions using an axiomatic variational approach. The proposed method allows to incorporate various generic region appearance models, while avoiding metrication errors. In the suggested framework, the segmentation is performed by level set evolution. Yet, contrarily to most existing methods, here, multiple regions are represented by a single non-negative level set function. The level set function evolution is efficiently executed through the Voronoi Implicit Interface Method for multi-phase interface evolution. The proposed approach is shown to obtain accurate segmentation results for various natural 2D and 3D images, comparable to state-of-the-art image segmentation algorithms. Anastasia Dubrovina, Guy Rosman, Ron Kimmel |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Aerial Reconstructions via Probabilistic Data FusionabstractWe propose an integrated probabilistic model for multi-modal fusion of aerial imagery, LiDAR data, and (optional) GPS measurements. The model allows for analysis and dense reconstruction (in terms of both geometry and appearance) of large 3D scenes. An advantage of the approach is that it explicitly models uncertainty and allows for missing data. As compared with image-based methods, dense reconstructions of complex urban scenes are feasible with fewer observations. Moreover, the proposed model allows one to estimate absolute scale and orientation and reason about other aspects of the scene, e.g., detection of moving objects. As formulated, the model lends itself to massively-parallel computing. We exploit this in an efficient inference scheme that utilizes both general purpose and domain-specific hardware components. We demonstrate results on large-scale reconstruction of urban terrain from LiDAR and aerial photography data. Randi Cabezas, Oren Freifeld, Guy Rosman, John W. Fisher III |
CVPR | 3 |
| 2014 | A Mixture of Manhattan Frames: Beyond the Manhattan WorldabstractObjects and structures within man-made environments typically exhibit a high degree of organization in the form of orthogonal and parallel planes. Traditional approaches to scene representation exploit this phenomenon via the somewhat restrictive assumption that every plane is perpendicular to one of the axes of a single coordinate system. Known as the Manhattan-World model, this assumption is widely used in computer vision and robotics. The complexity of many real-world scenes, however, necessitates a more flexible model. We propose a novel probabilistic model that describes the world as a mixture of Manhattan frames: each frame defines a different orthogonal coordinate system. This results in a more expressive model that still exploits the orthogonality constraints. We propose an adaptive Markov-Chain Monte-Carlo sampling algorithm with Metropolis-Hastings split/merge moves that utilizes the geometry of the unit sphere. We demonstrate the versatility of our Mixture-of-Manhattan-Frames model by describing complex scenes using depth images of indoor scenes as well as aerial-LiDAR measurements of an urban center. Additionally, we show that the model lends itself to focal-length calibration of depth cameras and to plane segmentation. Julian Straub, Guy Rosman, Oren Freifeld, John J. Leonard, John W. Fisher III |
CVPR | 2 |
| 2014 | Coresets for k-Segmentation of Streaming Data
Guy Rosman, Mikhail Volkov 0002, Dan Feldman, John W. Fisher III, Daniela Rus |
NIPS | 1 |
| 2013 | Patch-Collaborative Spectral Point-Cloud DenoisingabstractAbstract We present a new framework for point cloud denoising by patch‐collaborative spectral analysis. A collaborative generalization of each surface patch is defined, combining similar patches from the denoised surface. The Laplace–Beltrami operator of the collaborative patch is then used to selectively smooth the surface in a robust manner that can gracefully handle high levels of noise, yet preserves sharp surface features. The resulting denoising algorithm competes favourably with state‐of‐the‐art approaches, and extends patch‐based algorithms from the image processing domain to point clouds of arbitrary sampling. We demonstrate the accuracy and noise‐robustness of the proposed algorithm on standard benchmark models as well as range scans, and compare it to existing methods for point cloud denoising. Guy Rosman, Anastasia Dubrovina, Ron Kimmel |
Comput. Graph. Forum | 1 |
| 2012 | Fast Regularization of Matrix-Valued Images
Guy Rosman, Yu Wang 0029, Xue-Cheng Tai, Ron Kimmel, Alfred M. Bruckstein |
ECCV (3) | 1 |
| 2010 | Nonlinear Dimensionality Reduction by Topologically Constrained Isometric Embedding
Guy Rosman, Michael M. Bronstein, Alexander M. Bronstein, Ron Kimmel |
Int. J. Comput. Vis. | 1 |
| 2009 | Efficient Beltrami Image Filtering via Vector Extrapolation MethodsabstractThe Beltrami image flow is an effective nonlinear filter, often used in color image processing. It was shown to be closely related to the median, total variation, and bilateral filters. It treats the image as a two-dimensional manifold embedded in a hybrid spatial-feature space. Minimization of the image surface area yields the Beltrami flow. The corresponding diffusion operator is anisotropic and strongly couples the spectral components. Thus, there is so far no implicit or operator–splitting-based numerical scheme for the partial differential equation that describes the Beltrami flow in color. Usually, this flow is implemented by explicit schemes, which are stable only for very small time steps and therefore require many iterations. At the other end, vector extrapolation techniques accelerate the convergence of vector sequences, without explicit knowledge of the sequence generator. In this paper, we propose using vector extrapolation techniques for accelerating the convergence of the explicit schemes for the Beltrami flow. Experiments demonstrate fast convergence and efficiency compared to explicit schemes. Guy Rosman, Lorina Dascal, Avram Sidi, Ron Kimmel |
SIAM J. Imaging Sci. | 1 |
| 2007 | A New Physically Motivated Warping Model for Form Drop-OutabstractDocuments scanned by sheet-fed scanners often exhibit distortions due to the feeding and scanning mechanism. This paper presents a new model, motivated by the distortions observed in such documents. Numerical problems affecting the use of this model are addressed using an approximated model which is easier to estimate correctly. We demonstrate results showing the robustness and accuracy of this model on sheet-fed scanners output, and relate to existing techniques for registration and drop-out of structured forms. Guy Rosman, Asaf Tzadok, Doron Tal |
ICDAR | 1 |