VLDB 2026 Research / reviewers in the wild / expert
Masayoshi Tomizuka
dblp:10/4434
· DBLP profile ↗
220ranked-venue papers
0as first author
119since 2021 · last 2025
0000-0003-0206-6639ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 193 · 104 since 2021Systems, architecture and hardware · 117 · 55 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 13 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CompGS: Unleashing 2D Compositionality for Compositional Text-to-3D via Dynamically Optimizing 3D GaussiansabstractRecent breakthroughs in text-guided image generation have significantly advanced the field of 3D generation. While generating a single high-quality 3D object is now feasible, generating multiple objects with reasonable interactions within a 3D space, a.k.a. compositional 3D generation, presents substantial challenges. This paper introduces COMPGS, a novel generative framework that employs 3D Gaussian Splatting (GS) for efficient, compositional text-to-3D content generation. To achieve this goal, two core designs are proposed: (1) 3D Gaussians Initialization with 2D compositionality: We transfer the well-established 2D compositionality to initialize the Gaussian parameters on an entity-by-entity basis, ensuring both consistent 3D priors for each entity and reasonable interactions among multiple entities; (2) Dynamic Optimization: We propose a dynamic strategy to optimize 3D Gaussians using Score Distillation Sampling (SDS) loss. COMPGS first automatically decomposes 3D Gaussians into distinct entity parts, enabling optimization at both the entity and composition levels. Additionally, COMPGS optimizes across objects of varying scales by dynamically adjusting the spatial parameters of each entity, enhancing the generation of fine-grained details, particularly in smaller entities. Qualitative comparisons and quantitative evaluations on T3Bench demonstrate the effectiveness of COMPGS in generating compositional 3D objects with superior image quality and semantic alignment over existing methods. COMPGS can also be easily extended to progressively adding 3D objects, facilitating complex scene generation. We hope COMPGS will provide new insights to the compositional 3D generation. Chongjian Ge, Chenfeng Xu, Yuanfeng Ji, Chensheng Peng, Masayoshi Tomizuka, Ping Luo 0002, Mingyu Ding, Varun Jampani |
CVPR | 5 |
| 2025 | DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous ManipulationabstractDexterous manipulation with contact-rich interactions is crucial for advanced robotics. While recent diffusion-based planning approaches show promise for simple manipulation tasks, they often produce unrealistic ghost states (e.g., the object automatically moves without hand contact) or lack adaptability when handling complex sequential interactions. In this work, we introduce DexHand-Diff, an interaction-aware diffusion planning framework for adaptive dexterous manipulation. DexHandDiff models joint state-action dynamics through a dual-phase diffusion process which consists of pre-interaction contact alignment and post-contact goal-directed control, enabling goal-adaptive generalizable dexterous manipulation. Additionally, we incorporate dynamics model-based dual guidance and leverage large language models for automated guidance function generation, enhancing generalizability for physical interactions and facilitating diverse goal adaptation through language cues. Experiments on physical interaction tasks such as door opening, pen and block reorientation, object relocation, and hammer striking demonstrate DexHandDiff’s effectiveness on goals outside training distributions, achieving over twice the average success rate (59.2% vs. 29.5%) compared to existing methods. Our framework achieves an average of 70.7% success rate on goal adaptive dexterous tasks, highlighting its robustness and flexibility in contact-rich manipulation. Zhixuan Liang, Yao Mu 0001, Tianxing Chen, Wenqi Shao, Masayoshi Tomizuka, Ping Luo 0002, Mingyu Ding |
CVPR | 7 |
| 2025 | DeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving ScenesabstractWe present DeSiRe-GS, a self-supervised gaussian splatting representation, enabling effective static-dynamic decomposition and high-fidelity surface reconstruction in complex driving scenarios. Our approach employs a two-stage optimization pipeline of dynamic street Gaussians. In the first stage, we extract 2D motion masks based on the observation that 3D Gaussian Splatting inherently can reconstruct only the static regions in dynamic environments. These extracted 2D motion priors are then mapped into the Gaussian space in a differentiable manner, leveraging an efficient formulation of dynamic Gaussians in the second stage. Combined with the introduced geometric regularizations, our method are able to address the over-fitting issues caused by data sparsity in autonomous driving, reconstructing physically plausible Gaussians that align with object surfaces rather than floating in air. Furthermore, we introduce temporal cross-view consistency to ensure coherence across time and viewpoints, resulting in high-quality surface reconstruction. Comprehensive experiments demonstrate the efficiency and effectiveness of DeSiRe-GS, surpassing prior self-supervised arts and achieving accuracy comparable to methods relying on external 3D bounding box annotations. Code is available at https://github.com/chengweialan/DeSiRe-GS Chensheng Peng, Chenfeng Xu, Yichen Xie 0002, Wenzhao Zheng, Kurt Keutzer, Masayoshi Tomizuka |
CVPR | 8 |
| 2025 | Bridging Viewpoint Gaps: Geometric Reasoning Boosts Semantic CorrespondenceabstractFinding semantic correspondences between images is a challenging problem in computer vision, particularly under significant viewpoint changes. Previous methods rely on semantic features from pre-trained 2D models like Stable Diffusion and DINOv2, which often struggle to extract viewpoint-invariant features. To overcome this, we propose a novel approach that integrates geometric and semantic reasoning. Unlike prior methods relying on heuristic geometric enhancements, our framework fine-tunes DUSt3R on synthetic cross-instance data to reconstruct distinct objects in an aligned 3D space. By learning to deform these objects into similar shapes using semantic supervision, we enable efficient KNN-based geometric matching, followed by sparse semantic matching within local KNN candidates. While trained on synthetic data, our method generalizes effectively to real-world images, achieving up to 7.4-point improvements in zero-shot settings on the rigid-body subset of Spair-71K and up to 19.6-point gains under extreme viewpoint variations. Additionally, it accelerates runtime by up to 40 times, demonstrating both its robustness to viewpoint changes and its efficiency for practical applications. Code is available at https://github.com/QiyangQ/Bridge-Viewpoint-Gap. Qiyang Qian, Hansheng Chen 0004, Masayoshi Tomizuka, Kurt Keutzer, Qianqian Wang 0002, Chenfeng Xu |
CVPR | 3 |
| 2025 | LangTraj: Diffusion Model and Dataset for Language-Conditioned Trajectory SimulationabstractEvaluating autonomous vehicles with controllability enables scalable testing in counterfactual or structured settings, enhancing both efficiency and safety. We introduce LangTraj, a language-conditioned scene-diffusion model that simulates the joint behavior of all agents in traffic scenarios. By conditioning on natural language inputs, LangTraj provides flexible and intuitive control over interactive behaviors, generating nuanced and realistic scenarios. Unlike prior approaches that depend on domain-specific guidance functions, LangTraj incorporates language conditioning during training, facilitating more intuitive traffic simulation control. We propose a novel closed-loop training strategy for diffusion models, explicitly tailored to enhance stability and realism during closed-loop simulation. To support language-conditioned simulation, we develop Inter-Drive, a large-scale dataset with diverse and interactive labels for training language-conditioned diffusion models. Our dataset is built upon a scalable pipeline for annotating agent-agent interactions and single-agent behaviors, ensuring rich and varied supervision. Validated on the Waymo Open Motion Dataset, LangTraj demonstrates strong performance in realism, language controllability, and language-conditioned safety-critical simulation, establishing a new paradigm for flexible and scalable autonomous vehicle testing. Project Website: https://langtraj.github.io/ Wei-Jer Chang, Masayoshi Tomizuka, Manmohan Krishna Chandraker, Francesco Pittaluga |
ICCV | 3 |
| 2025 | StreamDiffusion: A Pipeline-Level Solution for Real-Time Interactive GenerationabstractWe introduce StreamDiffusion, a real-time diffusion pipeline designed for interactive image generation. Existing diffusion models are adept at creating images from text or image prompts, yet they often fall short in real-time interaction. This limitation becomes particularly evident in scenarios involving continuous input, such as Metaverse, live video streaming, and broadcasting, where high throughput is imperative. To address this, we present a novel approach that transforms the original sequential denoising into the batching denoising process. Stream Batch eliminates the conventional wait-and-interact approach and enables fluid and high throughput streams. To handle the frequency disparity between data input and model throughput, we design a novel input-output queue for parallelizing the streaming process. Moreover, the existing diffusion pipeline uses classifier-free guidance(CFG), which requires additional U-Net computation. To mitigate the redundant computations, we propose a novel residual classifier-free guidance (RCFG) algorithm that reduces the number of negative conditional denoising steps to only one or even zero. Besides, we introduce a stochastic similarity filter(SSF) to optimize power consumption. Our Stream Batch achieves around 1.5x speedup compared to the sequential denoising method at different denoising levels. The proposed RCFG leads to speeds up to 2.05x higher than the conventional CFG. Combining the proposed strategies and existing mature acceleration tools makes the image-to-image generation achieve up-to 91.07fps on one RTX4090, improving the throughputs of AutoPipline developed by Diffusers over 59.56x. Furthermore, our proposed StreamDiffusion also significantly reduces the energy consumption by 2.39x on one RTX3060 and 1.99x on one RTX4090, respectively. Akio Kodaira, Chenfeng Xu, Toshiki Hazama, Takanori Yoshimoto, Kohei Ohno, Shogo Mitsuhori, Soichi Sugano, Hanying Cho, Masayoshi Tomizuka, Kurt Keutzer |
ICCV | 10 |
| 2025 | A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D SupervisionabstractWe present a novel framework for training 3D image-conditioned diffusion models using only 2D supervision. Recovering 3D structure from 2D images is inherently ill-posed due to the ambiguity of possible reconstructions, making generative models a natural choice. However, most existing 3D generative models rely on full 3D supervision, which is impractical due to the scarcity of large-scale 3D datasets. To address this, we propose leveraging sparse-view supervision as a scalable alternative. While recent reconstruction models use sparse-view supervision with differentiable rendering to lift 2D images to 3D, they are predominantly deterministic, failing to capture the diverse set of plausible solutions and producing blurry predictions in uncertain regions. A key challenge in training 3D diffusion models with 2D supervision is that the standard training paradigm requires both the denoising process and supervision to be in the same modality. We address this by decoupling the noisy samples being denoised from the supervision signal, allowing the former to remain in 3D while the latter is provided in 2D. Our approach leverages suboptimal predictions from a deterministic image-to-3D model-acting as a "teacher"-to generate noisy 3D inputs, enabling effective 3D diffusion training without requiring full 3D ground truth. We validate our framework on both object-level and scene-level datasets, using two different 3D Gaussian Splat (3DGS) teachers. Our results show that our approach consistently improves upon these deterministic teachers, demonstrating its effectiveness in scalable and high-fidelity 3D generative modeling. See our project page at https://lesson-in-splats.github.io/ Chensheng Peng, Ido Sobol, Masayoshi Tomizuka, Kurt Keutzer, Chenfeng Xu, Or Litany |
ICCV | 3 |
| 2025 | X-Drive: Cross-modality Consistent Multi-Sensor Data Synthesis for Driving ScenariosabstractRecent advancements have exploited diffusion models for the synthesis of either LiDAR point clouds or camera image data in driving scenarios. Despite their success in modeling single-modality data marginal distribution, there is an under- exploration in the mutual reliance between different modalities to describe com- plex driving scenes. To fill in this gap, we propose a novel framework, X-DRIVE, to model the joint distribution of point clouds and multi-view images via a dual- branch latent diffusion model architecture. Considering the distinct geometrical spaces of the two modalities, X-DRIVE conditions the synthesis of each modality on the corresponding local regions from the other modality, ensuring better alignment and realism. To further handle the spatial ambiguity during denoising, we design the cross-modality condition module based on epipolar lines to adaptively learn the cross-modality local correspondence. Besides, X-DRIVE allows for controllable generation through multi-level input conditions, including text, bounding box, image, and point clouds. Extensive results demonstrate the high-fidelity synthetic results of X-DRIVE for both point clouds and multi-view images, adhering to input conditions while ensuring reliable cross-modality consistency. Our code will be made publicly available at https://github.com/yichen928/X-Drive. Yichen Xie 0002, Chenfeng Xu, Chensheng Peng, Shuqi Zhao, Nhat Ho, Alexander T. Pham, Mingyu Ding, Masayoshi Tomizuka |
ICLR | 8 |
| 2025 | Looking Backward: Streaming Video-to-Video Translation with Feature BanksabstractThis paper introduces StreamV2V, a diffusion model that achieves real-time streaming video-to-video (V2V) translation with user prompts.
Unlike prior V2V methods using batches to process limited frames, we opt to process frames in a streaming fashion, to support unlimited frames.
At the heart of StreamV2V lies a backward-looking principle that relates the present to the past.
This is realized by maintaining a feature bank, which archives information from past frames.
For incoming frames, StreamV2V extends self-attention to include banked keys and values, and directly fuses similar past features into the output.
The feature bank is continually updated by merging stored and new features, making it compact yet informative.
StreamV2V stands out for its adaptability and efficiency, seamlessly integrating with image diffusion models without fine-tuning.
It can run 20 FPS on one A100 GPU, being 15$\times$, 46$\times$, 108$\times$, and 158$\times$ faster than FlowVid, CoDeF, Rerender, and TokenFlow, respectively.
Quantitative metrics and user studies confirm StreamV2V's exceptional ability to maintain temporal consistency. Akio Kodaira, Chenfeng Xu, Masayoshi Tomizuka, Kurt Keutzer, Diana Marculescu |
ICLR | 4 |
| 2025 | Bisimulation Metric for Model Predictive ControlabstractModel-based reinforcement learning (MBRL) has shown promise for improving sample efficiency and decision-making in complex environments. However, existing methods face challenges in training stability, robustness to noise, and computational efficiency. In this paper, we propose Bisimulation Metric for Model Predictive Control (BS-MPC), a novel approach that incorporates bisimulation metric loss in its objective function to directly optimize the encoder. This optimization enables the learned encoder to extract intrinsic information from the original state space while discarding irrelevant details. BS-MPC improves training stability, robustness against input noise, and computational efficiency by reducing training time. We evaluate BS-MPC on both continuous control and image-based tasks from the DeepMind Control Suite, demonstrating superior performance and robustness compared to state-of-the-art baseline methods. Yutaka Shimizu, Masayoshi Tomizuka |
ICLR | 2 |
| 2025 | Dobi-SVD: Differentiable SVD for LLM Compression and Some New PerspectivesabstractLarge language models (LLMs) have sparked a new wave of AI applications; however, their substantial computational costs and memory demands pose significant challenges to democratizing access to LLMs for a broader audience. Singular Value Decomposition (SVD), a technique studied for decades, offers a hardware-independent and flexibly tunable solution for LLM compression. In this paper, we present new directions using SVD: we first theoretically analyze the optimality of truncating weights and truncating activations, then we further identify three key issues on SVD-based LLM compression, including (1) How can we determine the optimal truncation position for each weight matrix in LLMs? (2) How can we efficiently update the weight matrices based on truncation position? (3) How can we address the inherent "injection" nature that results in the information loss of the SVD? We propose an effective approach, **Dobi-SVD**, to tackle the three issues.
First, we propose a **differentiable** truncation-value learning mechanism, along with gradient-robust backpropagation, enabling the model to adaptively find the optimal truncation positions. Next, we utilize the Eckart-Young-Mirsky theorem to derive a theoretically **optimal** weight update formula through rigorous mathematical analysis. Lastly, by observing and leveraging the quantization-friendly nature of matrices after SVD decomposition, we reconstruct a mapping between truncation positions and memory requirements, establishing a **bijection** from truncation positions to memory.
Experimental results show that with a 40\% parameter-compression rate, our method achieves a perplexity of 9.07 on the Wikitext2 dataset with the compressed LLama-7B model, a 78.7\% improvement over the state-of-the-art SVD for LLM compression method.
We emphasize that Dobi-SVD is the first to achieve such a high-ratio LLM compression with minimal performance drop. We also extend our Dobi-SVD to VLM compression, achieving a 20\% increase in throughput with minimal performance degradation. We hope that the inference speedup—up to 12.4x on 12GB NVIDIA Titan Xp GPUs and 3x on 80GB A100 GPUs for LLMs, and 1.2x on 80GB A100 GPUs for VLMs—will bring significant benefits to the broader community such as robotics. Qinsi Wang, Jinghan Ke, Masayoshi Tomizuka, Kurt Keutzer, Chenfeng Xu |
ICLR | 3 |
| 2025 | Residual-MPPI: Online Policy Customization for Continuous ControlabstractPolicies developed through Reinforcement Learning (RL) and Imitation Learning (IL) have shown great potential in continuous control tasks, but real-world applications often require adapting trained policies to unforeseen requirements. While fine-tuning can address such needs, it typically requires additional data and access to the original training metrics and parameters.
In contrast, an online planning algorithm, if capable of meeting the additional requirements, can eliminate the necessity for extensive training phases and customize the policy without knowledge of the original training scheme or task. In this work, we propose a generic online planning algorithm for customizing continuous-control policies at the execution time, which we call Residual-MPPI. It can customize a given prior policy on new performance metrics in few-shot and even zero-shot online settings, given access to the prior action distribution alone. Through our experiments, we demonstrate that the proposed Residual-MPPI algorithm can accomplish the few-shot/zero-shot online policy customization task effectively, including customizing the champion-level racing agent, Gran Turismo Sophy (GT Sophy) 1.0, in the challenging car racing scenario, Gran Turismo Sport (GTS) environment. Code for MuJoCo experiments is included in the supplementary and will be open-sourced upon acceptance. Demo videos are available on our website: https://sites.google.com/view/residual-mppi. Pengcheng Wang 0004, Chenran Li, Catherine Weaver, Kenta Kawamoto, Masayoshi Tomizuka, Chen Tang 0001 |
ICLR | 5 |
| 2025 | WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in DrivingabstractLanguage models uncover unprecedented abilities in analyzing driving scenarios, owing to their limitless knowledge accumulated from text-based pre-training. Naturally, they should particularly excel in analyzing rule-based interactions, such as those triggered by traffic laws, which are well documented in texts. However, such interaction analysis remains underexplored due to the lack of dedicated language datasets that address it. Therefore, we propose Waymo Open Motion Dataset-Reasoning (WOMD-Reasoning), a comprehensive large-scale Q&As dataset built on WOMD focusing on describing and reasoning traffic rule-induced interactions in driving scenarios. WOMD-Reasoning also presents by far the largest multi-modal Q&A dataset, with 3 million Q&As on real-world driving scenarios, covering a wide range of driving topics from map descriptions and motion status descriptions to narratives and analyses of agents' interactions, behaviors, and intentions. To showcase the applications of WOMD-Reasoning, we design Motion-LLaVA, a motion-language model fine-tuned on WOMD-Reasoning. Quantitative and qualitative evaluations are performed on WOMD-Reasoning dataset as well as the outputs of Motion-LLaVA, supporting the data quality and wide applications of WOMD-Reasoning, in interaction predictions, traffic rule compliance plannings, etc. The dataset and its vision modal extension are available on https://waymo.com/open/download/. The codes & prompts to build it are available on https://github.com/yhli123/WOMD-Reasoning. Cunxin Fan, Chongjian Ge, Seth Z. Zhao, Chenran Li, Chenfeng Xu, Huaxiu Yao, Masayoshi Tomizuka, Bolei Zhou, Chen Tang 0001, Mingyu Ding |
ICML | 8 |
| 2025 | FDPP: Fine-Tune Diffusion Policy with Human PreferenceabstractImitation learning from human demonstrations enables robots to perform complex manipulation tasks and has recently witnessed huge success. However, these techniques often struggle to adapt behavior to new preferences or changes in the environment. To address these limitations, we propose Fine-tuning Diffusion Policy with Human Preference (FDPP). FDPP learns a reward function through preference-based learning. This reward is then used to fine-tune the pre-trained policy with reinforcement learning (RL), resulting in alignment of pre-trained policy with new human preferences while still solving the original task. Our experiments across various robotic tasks and preferences demonstrate that FDPP effectively customizes policy behavior without compromising performance. Additionally, we show that incorporating Kullback-Leibler (KL) regularization during fine-tuning prevents over-fitting and helps maintain the competencies of the initial policy. Devesh K. Jha, Masayoshi Tomizuka, Diego Romeres |
ICRA | 3 |
| 2025 | TrajSSL: Trajectory-Enhanced Semi-Supervised 3D Object DetectionabstractSemi-supervised 3D object detection is a common strategy employed to circumvent the challenge of manually labeling large-scale autonomous driving perception datasets. Pseudo-labeling approaches to semi-supervised learning adopt a teacher-student framework in which machine-generated pseudo-labels on a large unlabeled dataset are used in combination with a small manually-labeled dataset for training. In this work, we address the problem of improving pseudo-label quality through leveraging long- term temporal information captured in driving scenes. More specifically, we leverage pre-trained motion-forecasting models to generate object trajectories on pseudo-labeled data to further enhance the student model training. Our approach improves pseudo-label quality in two distinct manners: first, we suppress false positive pseudo-labels through establishing consistency across multiple frames of motion forecasting outputs. Second, we compensate for false negative detections by directly inserting predicted object tracks into the pseudo-labeled scene. Experiments on the nuScenes dataset demonstrate the effectiveness of our approach, improving the performance of standard semi-supervised approaches in a variety of settings. Philip L. Jacobson, Yichen Xie 0002, Mingyu Ding, Chenfeng Xu, Masayoshi Tomizuka, Ming C. Wu |
ICRA | 5 |
| 2025 | Adaptive Energy Regularization for Autonomous Gait Transition and Energy-Efficient Quadruped LocomotionabstractIn reinforcement learning for legged robot locomotion, crafting effective reward strategies is crucial. Predefined gait patterns and complex reward systems are widely used to stabilize policy training. Drawing from the natural locomotion behaviors of humans and animals, which adapt their gaits to minimize energy consumption, we investigate the impact of incorporating an energy-efficient reward term that prioritizes distance-averaged energy consumption into the reinforcement learning framework. Our findings demonstrate that this simple addition enables quadruped robots to autonomously select appropriate gaits-such as four-beat walking at lower speeds and trotting at higher speeds-without the need for explicit gait regularizations. Furthermore, we provide a guideline for tuning the weight of this energy-efficient reward, facilitating its application in real-world scenarios. The effectiveness of our approach is validated through simulations and on a real Unitree Gol robot. This research highlights the potential of energy-centric reward functions to simplify and enhance the learning of adaptive and efficient locomotion in quadruped robots. Videos and more details are at https://sites.google.com/berkeley.edu/efficient-locomotion Boyuan Liang, Lingfeng Sun, Xinghao Zhu, Bike Zhang, Ziyin Xiong, Chenran Li, Koushil Sreenath, Masayoshi Tomizuka |
ICRA | 9 |
| 2025 | Embodiment-agnostic Action Planning via Object-Part Scene FlowabstractObserving that the key for robotic action planning is to understand the target-object motion when its associated part is manipulated by the end effector, we propose to generate the 3D object-part scene flow and extract its transformations to solve the action trajectories for diverse embodiments. The advantage of our approach is that it derives the robot action explicitly from object motion prediction, yielding a more robust policy by understanding the object motions. Also, beyond policies trained on embodiment-centric data, our method is embodiment-agnostic, generalizable across diverse embodiments, and being able to learn from human demonstrations. Our method comprises three components: an object-part predictor to locate the part for the end effector to manipulate, an RGBD video generator to predict future RGBD videos, and a trajectory planner to extract embodiment-agnostic transformation sequences and solve the trajectory for diverse embodiments. Trained on videos even without trajectory data, our method still outperforms existing works significantly by 27.7% and 26.2% on the prevailing virtual environments MetaWorld and Franka-Kitchen, respectively. Furthermore, we conducted real-world experiments, showing that our policy, trained only with human demonstration, can be deployed to various embodiments. Weiliang Tang, Jia-Hui Pan, Jianshu Zhou, Huaxiu Yao, Yun-Hui Liu 0001, Masayoshi Tomizuka, Mingyu Ding, Chi-Wing Fu |
ICRA | 7 |
| 2025 | Cohere3D: Exploiting Temporal Coherence for Unsupervised Representation Learning of Vision-Based Autonomous DrivingabstractMulti-frame temporal inputs are important for vision-based autonomous driving. Observations from different angles enable the recovery of 3 D object states from 2 D images as long as we can identify the same instance from different input frames. However, the dynamic nature of driving scenes leads to significant variance in the instance appearance and shape captured by the cameras at different time steps. To this end, we propose a novel contrastive learning algorithm, Cohere3D, to learn coherent instance representations robust to the changes of distance and perspective in a long-term temporal sequence without any human annotations. In the pretraining stage, raw point clouds from LiDAR sensors are utilized to construct the instance-wise long-term temporal correspondence, which serves as guidance for the extraction of instance-level representation from the vision-based bird's-eye-view (BEV) feature map. Cohere3D encourages consistent representation for the same instance at different frames but distinguishes between different instances. We validate the effectiveness and generalizability of our algorithm by finetuning the pretrained model across key downstream autonomous driving tasks: perception, mapping, prediction, and planning. Results show a notable improvement in both data efficiency and final performance in all these tasks. Yichen Xie 0002, Hongge Chen, Gregory P. Meyer, Yong Jae Lee, Eric M. Wolff, Masayoshi Tomizuka, Yuning Chai |
ICRA | 6 |
| 2025 | Physics-Aware Robotic Palletization With Online Masking InferenceabstractThe efficient planning of stacking boxes, especially in the online setting where the sequence of item arrivals is unpredictable, remains a critical challenge in modern warehouse and logistics management. Existing solutions often address box size variations, but overlook their intrinsic and physical properties, such as density and rigidity, which are crucial for real-world applications. We use reinforcement learning (RL) to solve this problem by employing action space masking to direct the RL policy toward valid actions. Unlike previous methods that rely on heuristic stability assessments which are difficult to assess in physical scenarios, our framework utilizes online learning to dynamically train the action space mask, eliminating the need for manual heuristic design. Extensive experiments demonstrate that our proposed method outperforms existing state-of-the-arts. Furthermore, we deploy our learned task planner in a real-world robotic palletizer, validating its practical applicability in operational settings. The code is available at https://github.com/tianqi-zh/palletization. Zheng Wu 0002, Boyuan Liang, Scott Moura, Masayoshi Tomizuka, Mingyu Ding |
ICRA | 7 |
| 2025 | PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal ModelsabstractRobotic grasping, crucial for robot interaction with objects, still struggles with counter-intuitive or long-tailed scenarios like uncommon materials and shapes. Humans, however, intuitively adjust grasps with their physics-informed interpretations of the object, using visual and linguistic cues. This work introduces PhyGrasp, a large multimodal model and dataset that enhance robotic manipulation by combining natural language and 3D point clouds using a bridge module to integrate these inputs. The language modality exhibits robust reasoning capabilities concerning the impacts of diverse physical properties on grasping, while the 3D modality comprehends object shapes and parts. With these two capabilities, PhyGrasp is able to accurately assess the physical properties of object parts and determine optimal grasping poses. Additionally, the model’s language comprehension enables human instruction interpretation, generating grasping poses that align with human preferences. To train PhyGrasp, we construct a dataset PhyPartNet with 195K object instances with varying physical properties and human preferences, alongside their corresponding language descriptions. Extensive experiments conducted in the simulation and on the real robots demonstrate that PhyGrasp achieves state-of-the-art performance, particularly in long-tailed cases, e.g., about 10% improvement in success rate over GraspNet. More demos and information are available on https://sites.google.com/view/phygrasp. Dingkun Guo, Yuqi Xiang, Shuqi Zhao, Xinghao Zhu, Masayoshi Tomizuka, Mingyu Ding |
IROS | 5 |
| 2025 | P2 Explore: Efficient Exploration in Unknown Cluttered Environment with Floor Plan PredictionabstractRobot exploration aims at the reconstruction of unknown environments, and it is important to achieve it with shorter paths. Traditional methods focus on optimizing the visiting order of frontiers based on current observations, which may lead to local-minimal results. Recently, by predicting the structure of the unseen environment, the exploration efficiency can be further improved. However, in a cluttered environment, due to the randomness of obstacles, the ability to predict is weak. Moreover, this inaccuracy will lead to limited improvement in exploration. Therefore, we propose FPUNet which can be efficient in predicting the layout of noisy indoor environments. Then, we extract the segmentation of rooms and construct their topological connectivity based on the predicted map. The visiting order of these predicted rooms is optimized which can provide high-level guidance for exploration. The FPUNet is compared with other network architectures which demonstrates it is the SOTA method for this task. Extensive experiments in simulations show that our method can shorten the path length by 2.18% to 34.60% compared to the baselines. Gaoming Chen, Masayoshi Tomizuka, Zhenhua Xiong 0001, Mingyu Ding |
IROS | 3 |
| 2025 | DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous DrivingabstractLarge reconstruction model has remarkable progress, which can directly predict 3D or 4D representations for unseen scenes and objects. However, current work has not systematically explored the potential of large reconstruction models in the field of autonomous driving. To achieve this, we introduce the Large 4D Gaussian Reconstruction Model (DrivingRecon). With an elaborate and simple framework design, it not only ensures efficient and high-quality reconstruction, but also provides potential for downstream tasks. There are two core contributions: firstly, the Prune and Dilate Block (PD-Block) is proposed to prune redundant and overlapping Gaussian points and dilate Gaussian points for complex objects. Then, dynamic and static decoupling is tailored to better learn the temporary-consistent geometry across different time. Experimental results demonstrate that DrivingRecon significantly improves scene reconstruction quality compared to existing methods. Furthermore, we explore applications of DrivingRecon in model pre-training, vehicle type adaptation, and scene editing. Our code will be available. Hao Lu 0009, Tianshuo Xu, Wenzhao Zheng, Dalong Du, Masayoshi Tomizuka, Kurt Keutzer, Ying-Cong Chen |
NeurIPS | 7 |
| 2025 | Bootstrap Off-policy with World ModelabstractOnline planning has proven effective in reinforcement learning (RL) for improving sample efficiency and final performance. However, using planning for environment interaction inevitably introduces a divergence between the collected data and the policy's actual behaviors, degrading both model learning and policy improvement. To address this, we propose BOOM (Bootstrap Off-policy with WOrld Model), a framework that tightly integrates planning and off-policy learning through a bootstrap loop: the policy initializes the planner, and the planner refines actions to bootstrap the policy through behavior alignment. This loop is supported by a jointly learned world model, which enables the planner to simulate future trajectories and provides value targets to facilitate policy improvement. The core of BOOM is a likelihood-free alignment loss that bootstraps the policy using the planner’s non-parametric action distribution, combined with a soft value-weighted mechanism that prioritizes high-return behaviors and mitigates variability in the planner’s action quality within the replay buffer. Experiments on the high-dimensional DeepMind Control Suite and Humanoid-Bench show that BOOM achieves state-of-the-art results in both training stability and final performance. The code is accessible at \url{https://github.com/molumitu/BOOM_MBRL}. Guojian Zhan, Xiangteng Zhang, Jiaxin Gao 0002, Masayoshi Tomizuka, Shengbo Eben Li |
NeurIPS | 5 |
| 2025 | Sel4FT: Annotation Selection for Pretraining-Finetuning With Distribution ShiftabstractThe pretraining-finetuning paradigm has become dominant in computer vision, yet strategically exploiting limited annotation budgets during finetuning remains unexplored. We introduce active finetuning-a novel task for selecting the most informative samples to annotate within this paradigm. We propose Sel4FT, a unified annotation selection framework that optimizes a parametric model in continuous feature space to identify a subset preserving the entire pool's distribution while maintaining diversity. To address distribution shifts from data augmentation, we develop Sel4FT++ with augmentation-aware selection mechanisms. We theoretically prove our approach minimizes the Earth Mover's Distance between selected subset and full data pool. Our framework eliminates iterative retraining and annotation process during selection, providing an efficient solution for real-world deployment. Extensive experiments on image classification, long-tailed recognition, and semantic segmentation demonstrate state-of-the-art performance with over $100\times$100× speedup compared to existing methods. Han Lu 0004, Yichen Xie 0002, Mingyu Ding, Xiaokang Yang 0001, Masayoshi Tomizuka, Junchi Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Policy-Iteration-Based Active Disturbance Rejection Control for Uncertain Nonlinear Systems With Unknown Relative DegreeabstractIn this article, a policy-iteration-based active disturbance rejection control (ADRC) is proposed for uncertain nonlinear systems to achieve real-time output tracking performance, regardless of the specific relative degree of the system. The approach integrates a partial control input generator with a policy-iteration-based reinforcement learning (RL) agent for degree weight adjustment. The partial control input generator includes each ith order partial control input, which is constructed following the ADRC design framework for an ith order system. The RL agent adjusts the degree weights (its actions) to enhance the dominance of the partial control input corresponding to the unknown relative degree through iterative policy refinement. The RL agent is designed to minimize the quadratic reward as the performance index function while enhancing the influence of the partial control input associated with the correct relative degree via the policy iteration procedure. All signals in the closed-loop system (including the time-varying degree weights) ensure semi-global uniformly ultimately boundness using the Lyapunov stability theorem and the affinely quadratically stable property. Consequently, the degree weight adjustments by the RL agent do not affect the closed-loop stability. The proposed method does not require system dynamics, specific relative degree, external disturbances, and other state variable sensing beyond output sensing. The performance of the proposed method was validated via simulations for two different-order uncertain nonlinear systems and experiments using a permanent magnet synchronous motor testbed. Sesun You, Kwankyun Byeon, Masayoshi Tomizuka |
IEEE Trans. Cybern. | 5 |
| 2025 | Programmable Locking Cells (PLC) for Modular Robots With High Stiffness Tunability and Morphological Adaptability
Jianshu Zhou, Wei Chen 0068, Junda Huang, Boyuan Liang, Yun-Hui Liu 0001, Masayoshi Tomizuka |
IEEE Trans. Robotics | 6 |
| 2024 | Towards Generalizable and Interpretable Motion Prediction: A Deep Variational Bayes ApproachabstractEstimating the potential behavior of the surrounding human-driven vehicles is crucial for the safety of autonomous vehicles in a mixed traffic flow. Recent state-of-the-art achieved accurate prediction using deep neural networks. However, these end-to-end models are usually black boxes with weak interpretability and generalizability. This paper proposes the Goal-based Neural Variational Agent (GNeVA), an interpretable generative model for motion prediction with robust generalizability to out-of-distribution cases. For interpretability, the model achieves target-driven motion prediction by estimating the spatial distribution of long-term destinations with a variational mixture of Gaussians. We identify a causal structure among maps and agents’ histories and derive a variational posterior to enhance generalizability. Experiments on motion prediction datasets validate that the fitted model can be interpretable and generalizable and can achieve comparable performance to state-of-the-art results. Juanwu Lu, Masayoshi Tomizuka, Yeping Hu |
AISTATS | 3 |
| 2024 | SkillDiffuser: Interpretable Hierarchical Planning via Skill Abstractions in Diffusion-Based Task ExecutionabstractDiffusion models have demonstrated strong potential for robotic trajectory planning. However, generating coherent trajectories from high-level instructions remains challenging, especially for long-range composition tasks requiring multiple sequential skills. We propose SkillDiffuser, an end-to-end hierarchical planning framework integrating interpretable skill learning with conditional diffusion planning to address this problem. At the higher level, the skill abstraction module learns discrete, human-understandable skill representations from visual observations and language instructions. These learned skill embeddings are then used to condition the diffusion model to generate customized latent trajectories aligned with the skills. This allows generating diverse state trajectories that adhere to the learnable skills. By integrating skill learning with conditional trajectory generation, SkillDiffuser produces coherent behavior following abstract instructions across diverse tasks. Experiments on multitask robotic manipulation benchmarks like Meta-World and LOReL demonstrate state-of-the-art performance and human-interpretable skill representations from SkillDiffuser. More visualization results and information could be found on our website. Zhixuan Liang, Yao Mu 0001, Hengbo Ma, Masayoshi Tomizuka, Mingyu Ding, Ping Luo 0002 |
CVPR | 4 |
| 2024 | SAFE-SIM: Safety-Critical Closed-Loop Traffic Simulation with Diffusion-Controllable Adversaries
Wei-Jer Chang, Francesco Pittaluga, Masayoshi Tomizuka, Manmohan Krishna Chandraker |
ECCV (21) | 3 |
| 2024 | Optimizing Diffusion Models for Joint Trajectory Prediction and Controllable Generation
Chen Tang 0001, Lingfeng Sun, Simone Rossi 0001, Yichen Xie 0002, Chensheng Peng, Thomas Hannagan, Stefano Sabatini, Nicola Poerio, Masayoshi Tomizuka |
ECCV (29) | 10 |
| 2024 | UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingabstractLarge-scale vision-language pre-trained models have shown promising transferability to various downstream tasks. As the size of these foundation models and the number of downstream tasks grow, the standard full fine-tuning paradigm becomes unsustainable due to heavy computational and storage costs. This paper proposes UniAdapter, which unifies unimodal and multimodal adapters for parameter-efficient cross-modal adaptation on pre-trained vision-language models. Specifically, adapters are distributed to different modalities and their interactions, with the total number of tunable parameters reduced by partial weight sharing. The unified and knowledge-sharing design enables powerful cross-modal representations that can benefit various downstream tasks, requiring only 1.0%-2.0% tunable parameters of the pre-trained model. Extensive experiments on 7 cross-modal downstream benchmarks (including video-text retrieval, image-text retrieval, VideoQA, VQA and Caption) show that in most cases, UniAdapter not only outperforms the state-of-the-arts, but even beats the full fine-tuning strategy. Particularly, on the MSRVTT retrieval task, UniAdapter achieves 49.7% recall@1 with 2.2% model parameters, outperforming the latest competitors by 2.0%. The code and models are available at https://github.com/RERV/UniAdapter. Haoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu 0001, Masayoshi Tomizuka, Mingyu Ding |
ICLR | 6 |
| 2024 | What Matters to You? Towards Visual Representation Alignment for Robot LearningabstractWhen operating in service of people, robots need to optimize rewards aligned with end-user preferences. Since robots will rely on raw perceptual inputs, their rewards will inevitably use visual representations. Recently there has been excitement in using representations from pre-trained visual models, but key to making these work in robotics is fine-tuning, which is typically done via proxy tasks like dynamics prediction or enforcing temporal cycle-consistency. However, all these proxy tasks bypass the human’s input on what matters to them, exacerbating spurious correlations and ultimately leading to behaviors that are misaligned with user preferences. In this work, we propose that robots should leverage human feedback to align their visual representations with the end-user and disentangle what matters for the task. We propose Representation-Aligned Preference-based Learning (RAPL), a method for solving the visual representation alignment problem and visual reward learning problem through the lens of preference-based learning and optimal transport. Across experiments in X MAGICAL and in robotic manipulation, we find that RAPL’s reward consistently generates preferred robot behaviors with high sample efficiency, and shows strong zero-shot generalization when the visual representation is learned from a different embodiment than the robot’s. Thomas Tian, Chenfeng Xu, Masayoshi Tomizuka, Jitendra Malik, Andrea Bajcsy |
ICLR | 3 |
| 2024 | Guided Online Distillation: Promoting Safe Reinforcement Learning by Offline DemonstrationabstractSafe Reinforcement Learning (RL) aims to find a policy that achieves high rewards while satisfying cost constraints. When learning from scratch, safe RL agents tend to be overly conservative, which impedes exploration and restrains the overall performance. In many realistic tasks, e.g. autonomous driving, large-scale expert demonstration data are available. We argue that extracting expert policy from offline data to guide online exploration is a promising solution to mitigate the conserveness issue. Large-capacity models, e.g. decision transformers (DT), have been proven to be competent in offline policy learning. However, data collected in realworld scenarios rarely contain dangerous cases (e.g., collisions), which makes it prohibitive for the policies to learn safety concepts. Besides, these bulk policy networks cannot meet the computation speed requirements at inference time on real-world tasks such as autonomous driving. To this end, we propose Guided Online Distillation (GOLD), an offline-to-online safe RL framework. GOLD distills an offline DT policy into a lightweight policy network through guided online safe RL training, which outperforms both the offline DT policy and online safe RL algorithms. Experiments in both benchmark safe RL tasks and real-world driving tasks based on the Waymo Open Motion Dataset (WOMD) [1] demonstrate that GOLD can successfully distill lightweight policies and solve decision-making problems in challenging safety-critical scenarios. Jinning Li 0002, Banghua Zhu, Jiantao Jiao, Masayoshi Tomizuka, Chen Tang 0001 |
ICRA | 5 |
| 2024 | Robust In-Hand Manipulation with Extrinsic ContactsabstractWe present in-hand manipulation tasks where a robot moves an object in grasp, maintains its external contact mode with the environment, and adjusts its in-hand pose simultaneously. The proposed manipulation task leads to complex contact interactions which can be very susceptible to uncertainties in kinematic and physical parameters. Therefore, we propose a robust in-hand manipulation method, which consists of two parts. First, an in-gripper mechanics model that computes a naïve motion cone assuming all parameters are precise. Then, a robust planning method refines the motion cone to maintain desired contact mode regardless of parametric errors. Real-world experiments were conducted to illustrate the accuracy of the mechanics model and the effectiveness of the robust planning framework in the presence of kinematics parameter errors. Boyuan Liang, Kei Ota, Masayoshi Tomizuka, Devesh K. Jha |
ICRA | 3 |
| 2024 | Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment CollaborationabstractLarge, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io. Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin |
ICRA | 153 |
| 2024 | Interactive Planning Using Large Language Models for Partially Observable Robotic TasksabstractDesigning robotic agents to perform open vocabulary tasks has been the long-standing goal in robotics and AI. Recently, Large Language Models (LLMs) have achieved impressive results in creating robotic agents for performing open vocabulary tasks. However, planning for these tasks in the presence of uncertainties is challenging as it requires "chain-of-thought" reasoning, aggregating information from the environment, updating state estimates, and generating actions based on the updated state estimates. In this paper, we present an interactive planning technique for partially observable tasks using LLMs. In the proposed method, an LLM is used to collect missing information from the environment using a robot, and infer the state of the underlying problem from collected observations while guiding the robot to perform the required actions. We also use a fine-tuned Llama 2 model via self-instruct and compare its performance against a pre-trained LLM like GPT-4. Results are demonstrated on several tasks in simulation as well as real-world environments. Lingfeng Sun, Devesh K. Jha, Chiori Hori, Siddarth Jain, Radu Corcodel, Xinghao Zhu, Masayoshi Tomizuka, Diego Romeres |
ICRA | 7 |
| 2024 | MATRIX: Multi-Agent Trajectory Generation with Diverse ContextsabstractData-driven methods have great advantages in modeling complicated human behavioral dynamics and dealing with many human-robot interaction applications. However, collecting massive and annotated real-world human datasets has been a laborious task, especially for highly interactive scenarios. On the other hand, algorithmic data generation methods are usually limited by their model capacities, making them unable to offer realistic and diverse data needed by various application users. In this work, we study trajectory-level data generation for multi-human or human-robot interaction scenarios and propose a learning-based automatic trajectory generation model, which we call Multi-Agent TRajectory generation with dIverse conteXts (MATRIX). MATRIX is capable of generating interactive human behaviors in realistic diverse contexts. We achieve this goal by modeling the explicit and interpretable objectives so that MATRIX can generate human motions based on diverse destinations and heterogeneous behaviors. We carried out extensive comparison and ablation studies to illustrate the effectiveness of our approach across various metrics. We also presented experiments that demonstrate the capability of MATRIX to serve as data augmentation for imitation-based motion planning. Yida Yin, Huidong Gao, Masayoshi Tomizuka, Jiachen Li 0001 |
ICRA | 5 |
| 2024 | Bridging the Sim-to-Real Gap with Dynamic Compliance Tuning for Industrial InsertionabstractContact-rich manipulation tasks often exhibit a large sim-to-real gap. For instance, industrial assembly tasks frequently involve tight insertions where the clearance is less than 0.1 mm and can even be negative when dealing with a deformable receptacle. This narrow clearance leads to complex contact dynamics that are difficult to model accurately in simulation, making it challenging to transfer simulation-learned policies to real-world robots. In this paper, we propose a novel framework for robustly learning manipulation skills for real-world tasks using simulated data only. Our framework consists of two main components: the "Force Planner" and the "Gain Tuner". The Force Planner plans both the robot motion and desired contact force, while the Gain Tuner dynamically adjusts the compliance control gains to track the desired contact force during task execution. The key insight is that by dynamically adjusting the robot’s compliance control gains during task execution, we can modulate contact force in the new environment, thereby generating trajectories similar to those trained in simulation and narrowing the sim-to-real gap. Experimental results show that our method, trained in simulation on a generic square peg-and-hole task, can generalize to a variety of real-world insertion tasks involving narrow and negative clearances, all without requiring any fine-tuning. Videos are available at https://dynamic-compliance.github.io Xiang Zhang 0020, Masayoshi Tomizuka, Hui Li 0009 |
ICRA | 2 |
| 2024 | Multi-level Reasoning for Robotic Assembly: From Sequence Inference to Contact SelectionabstractAutomating the assembly of objects from their parts is a complex problem with innumerable applications in manufacturing, maintenance, and recycling. Unlike existing research, which is limited to target segmentation, pose regression, or using fixed target blueprints, our work presents a holistic multi-level framework for part assembly planning consisting of part assembly sequence inference, part motion planning, and robot contact optimization. We present the Part Assembly Sequence Transformer (PAST) – a sequence-to-sequence neural network – to infer assembly sequences recursively from a target blueprint. We then use a motion planner and optimization to generate part movements and contacts. To train PAST, we introduce D4PAS: a large-scale Dataset for Part Assembly Sequences consisting of physically valid sequences for industrial objects. Experimental results show that our approach generalizes better than prior methods while needing significantly less computational time for inference. Further details on our experiments and results are available in the video. Xinghao Zhu, Devesh K. Jha, Diego Romeres, Lingfeng Sun, Masayoshi Tomizuka, Anoop Cherian |
ICRA | 5 |
| 2024 | Joint Pedestrian Trajectory Prediction through Posterior SamplingabstractJoint pedestrian trajectory prediction has long grappled with the inherent unpredictability of human behaviors. Recent works employing conditional diffusion models in trajectory prediction have exhibited notable success. Nevertheless, the heavy dependence on accurate historical data results in their vulnerability to noise disturbances and data incompleteness. To improve the robustness and reliability, we introduce the Guided Full Trajectory Diffuser (GFTD), a novel diffusion-based framework that translates prediction as the inverse problem of spatial-temporal inpainting and models the full joint trajectory distribution which includes both history and the future. By learning from the full trajectory and leveraging flexible posterior sampling methods, GFTD can produce accurate predictions while improving the robustness that can generalize to scenarios with noise perturbation or incomplete historical data. Moreover, the pre-trained model enables controllable generation without an additional training budget. Through rigorous experimental evaluation, GFTD exhibits superior performance in joint trajectory prediction with different data quality and in controllable generation tasks. See more results at https://sites.google.com/andrew.cmu.edu/posterior-sampling-prediction. Haotian Lin 0003, Mingxiao Huo, Chensheng Peng, Masayoshi Tomizuka |
IROS | 6 |
| 2024 | Harnessing with Twisting: Single-Arm Deformable Linear Object Manipulation for Industrial Harnessing TaskabstractWire-harnessing tasks pose great challenges to be automated by the robot due to the complex dynamics and unpredictable behavior of the deformable wire. Traditional methods, often reliant on dual-robot arms or tactile sensing, face limitations in adaptability, cost, and scalability. This paper introduces a novel single-robot wire-harnessing pipeline that leverages a robot’s twisting motion to generate necessary wire tension for precise insertion into clamps, using only one robot arm with an integrated force/torque (F/T) sensor. Benefiting from this design, the single robot arm can efficiently apply tension for wire routing and insertion into clamps in a narrow space. Our approach is structured around four principal components: a Model Predictive Control (MPC) based on the Koopman operator for tension tracking and wire following, a motion planner for sequencing harnessing waypoints, a suite of insertion primitives for clamp engagement, and a fix-point switching mechanism for wire constraint updating. Evaluated on an industrial-level wire harnessing task, our method demonstrated superior performance and reliability over conventional approaches, efficiently handling both single and multiple wire configurations with high success rates. Xiang Zhang 0020, Hsien-Chung Lin, Yu Zhao 0015, Masayoshi Tomizuka |
IROS | 4 |
| 2024 | Contact-Implicit Model Predictive Control for Dexterous In-hand Manipulation: A Long-Horizon and Robust ApproachabstractDexterous in-hand manipulation is an essential skill of production and life. However, the highly stiff and mutable nature of contacts limits real-time contact detection and inference, degrading the performance of model-based methods. Inspired by recent advances in contact-rich locomotion and manipulation, this paper proposes a novel model-based approach to control dexterous in-hand manipulation and overcome the current limitations. The proposed approach has an attractive feature, which allows the robot to robustly perform long-horizon in-hand manipulation without predefined contact sequences or separate planning procedures. Specifically, we design a high-level contact-implicit model predictive controller to generate real-time contact plans executed by the low-level tracking controller. Compared to other model-based methods, such a long-horizon feature enables replanning and robust execution of contact-rich motions to achieve large displacements in-hand manipulation more efficiently; Compared to existing learning-based methods, the proposed approach achieves dexterity and also generalizes to different objects without any pre-training. Detailed simulations and ablation studies demonstrate the efficiency and effectiveness of our method. It runs at 20Hz on the 23-degree-of-freedom, long-horizon, in-hand object rotation task. Yongpeng Jiang, Mingrui Yu 0001, Xinghao Zhu, Masayoshi Tomizuka, Xiang Li 0009 |
IROS | 4 |
| 2024 | Pre-training on Synthetic Driving Data for Trajectory PredictionabstractAccumulating substantial volumes of real-world driving data proves pivotal in the realm of trajectory forecasting for autonomous driving. Given the heavy reliance of current trajectory forecasting models on data-driven methodologies, we aim to tackle the challenge of learning general trajectory forecasting representations under limited data availability. We propose a pipeline-level solution to mitigate the issue of data scarcity in trajectory forecasting. The solution is composed of two parts: firstly, we adopt HD map augmentation and trajectory synthesis for generating driving data, and then we learn representations by pre-training on them. Specifically, we apply vector transformations to reshape the maps, and then employ a rule-based model to generate trajectories on both original and augmented scenes; thus enlarging the driving data without collecting additional real ones. To foster the learning of general representations within this augmented dataset, we comprehensively explore the different pre-training strategies, including extending the concept of a Masked AutoEncoder (MAE) for trajectory forecasting. Without bells and whistles, our proposed pipeline-level solution is general, simple, yet effective: we conduct extensive experiments to demonstrate the effectiveness of our data expansion and pre-training strategies, which outperform the baseline prediction model by large margins, e.g. 5.04%, 3.84% and 8.30% in terms of MR6, minADE6and minFDE6. The pre-training dataset and the codes for pre-training and fine-tuning are released at https://github.com/yhli123/Pretraining_on_Synthetic_Driving_Data_for_Trajectory_Prediction. Seth Z. Zhao, Chenfeng Xu, Chen Tang 0001, Chenran Li, Mingyu Ding, Masayoshi Tomizuka |
IROS | 7 |
| 2024 | In-Hand Following of Deformable Linear Objects Using Dexterous Fingers with Tactile SensingabstractMost research on deformable linear object (DLO) manipulation assumes rigid grasping. However, beyond rigid grasping and re-grasping, in-hand following is also an essential skill that humans use to dexterously manipulate DLOs, which requires continuously changing the grasp point by in-hand sliding while holding the DLO to prevent it from falling. Achieving such a skill is very challenging for robots without using specially designed but not versatile end-effectors. Previous works have attempted using generic parallel grippers, but their robustness is unsatisfactory owing to the conflict between following and holding, which is hard to balance with a one-degree-of-freedom gripper. In this work, inspired by how humans use fingers to follow DLOs, we explore the usage of a generic dexterous hand with tactile sensing to imitate human skills and achieve robust in-hand DLO following. To enable the hardware system to function in the real world, we develop a framework that includes Cartesian-space arm-hand control, tactile-based in-hand 3-D DLO pose estimation, and task-specific motion design. Experimental results demonstrate the significant superiority of our method over using parallel grippers, as well as its great robustness, generalizability, and efficiency. Mingrui Yu 0001, Boyuan Liang, Xiang Zhang 0020, Xinghao Zhu, Lingfeng Sun, Shiji Song, Xiang Li 0009, Masayoshi Tomizuka |
IROS | 9 |
| 2024 | DSLO: Deep Sequence LiDAR Odometry Based on Inconsistent Spatio-temporal PropagationabstractThis paper introduces a 3D point cloud sequence learning model based on inconsistent spatio-temporal propagation for LiDAR odometry, termed DSLO. It consists of a pyramid structure with a spatial information reuse strategy, a sequential pose initialization module, a gated hierarchical pose refinement module, and a temporal feature propagation module. First, spatial features are encoded using a point feature pyramid, with features reused in successive pose estimations to reduce computational overhead. Second, a sequential pose initialization method is introduced, leveraging the high-frequency sampling characteristic of LiDAR to initialize the LiDAR pose. Then, a gated hierarchical pose refinement mechanism refines poses from coarse to fine by selectively retaining or discarding motion information from different layers based on gate estimations. Finally, temporal feature propagation is proposed to incorporate the historical motion information from point cloud sequences, and address the spatial inconsistency issue when transmitting motion information embedded in point clouds between frames. Experimental results on the KITTI odometry dataset and Argoverse dataset demonstrate that DSLO outperforms state-of-the-art methods, achieving at least a 15.67% improvement on RTE and a 12.64% improvement on RRE, while also achieving a 34.69% reduction in runtime compared to baseline methods. Our implementation will be available at https://github.com/IRMVLab/DSLO. Guangming Wang 0001, Xinrui Wu, Chenfeng Xu, Mingyu Ding, Masayoshi Tomizuka, Hesheng Wang 0001 |
IROS | 6 |
| 2024 | Immiscible Diffusion: Accelerating Diffusion Training with Noise AssignmentabstractIn this paper, we point out that suboptimal noise-data mapping leads to slow training of diffusion models. During diffusion training, current methods diffuse each image across the entire noise space, resulting in a mixture of all images at every point in the noise layer. We emphasize that this random mixture of noise-data mapping complicates the optimization of the denoising function in diffusion models. Drawing inspiration from the immiscibility phenomenon in physics, we propose *Immiscible Diffusion*, a simple and effective method to improve the random mixture of noise-data mapping. In physics, miscibility can vary according to various intermolecular forces. Thus, immiscibility means that the mixing of molecular sources is distinguishable. Inspired by this concept, we propose an assignment-then-diffusion training strategy to achieve *Immiscible Diffusion*. As one example, prior to diffusing the image data into noise, we assign diffusion target noise for the image data by minimizing the total image-noise pair distance in a mini-batch. The assignment functions analogously to external forces to expel the diffuse-able areas of images, thus mitigating the inherent difficulties in diffusion training. Our approach is remarkably simple, requiring only *one line of code* to restrict the diffuse-able area for each image while preserving the Gaussian distribution of noise. In this way, each image is preferably projected to nearby noise. To address the high complexity of the assignment algorithm, we employ a quantized assignment strategy, which significantly reduces the computational overhead to a negligible level (e.g. 22.8ms for a large batch size of 1024 on an A6000). Experiments demonstrate that our method can achieve up to 3x faster training for unconditional Consistency Models on the CIFAR dataset, as well as for DDIM and Stable Diffusion on CelebA and ImageNet dataset, and in class-conditional training and fine-tuning. In addition, we conducted a thorough analysis that sheds light on how it improves diffusion training speed while improving fidelity. The code is available at https://yhli123.github.io/immiscible-diffusion Heyang Jiang, Akio Kodaira, Masayoshi Tomizuka, Kurt Keutzer, Chenfeng Xu |
NeurIPS | 4 |
| 2024 | Data-Driven Linear Quadratic Optimization for Controller Synthesis With Structural ConstraintsabstractFor various typical cases and situations where the formulation results in an optimal control problem, the linear quadratic regulator (LQR) approach and its variants continue to be highly attractive. In certain scenarios, it can happen that some prescribed structural constraints on the gain matrix would arise. Consequently then, the algebraic Riccati equation (ARE) is no longer applicable in a straightforward way to obtain the optimal solution. This work presents a rather effective alternative optimization approach based on gradient projection. The utilized gradient is obtained through a data-driven methodology, and then projected onto applicable constrained hyperplanes. Essentially, this projection gradient determines a direction of progression and computation for the gain matrix update with a decreasing functional cost; and then the gain matrix is further refined in an iterative framework. With this formulation, a data-driven optimization algorithm is summarized for controller synthesis with structural constraints. This data-driven approach has the key advantage that it avoids the necessity of precise modeling which is always required in the classical model-based counterpart; and thus the approach can additionally accommodate various model uncertainties. Illustrative examples are also provided in the work to validate the theoretical results. Jun Ma 0008, Zilong Cheng, Xiaocong Li, Masayoshi Tomizuka, Tong Heng Lee |
IEEE Trans. Cybern. | 5 |
| 2024 | Real-Time Iterative Compensation Control Using Plant-Injection Feedforward Architecture With Application to Ultraprecision Wafer StagesabstractThis article presents a novel real-time iterative compensation (RIC) method using plant-injection feedforward architecture for motion control of ultraprecision wafer stages, addressing the severe challenge of achieving extreme tracking accuracy along with strong task flexibility and disturbance rejection ability. The RIC method establishes an online prediction model to accurately predict upcoming tracking errors during real-time motion. The prediction result enables the online generation of optimal plant-injection feedforward signal at each sampling control instant via iterative calculation, which enhances tracking accuracy and dynamical regulation capability. Various trajectory tracking tasks have been implemented on an ultraprecision wafer stage. Experimental results demonstrate that RIC matches the high tracking accuracy of well-acknowledged iterative learning control while offering superior task flexibility and disturbance rejection ability. Ran Zhou 0001, Chuxiong Hu, Ze Wang 0002, Yu Zhu 0001, Masayoshi Tomizuka |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | Grounded Relational Inference: Domain Knowledge Driven Explainable Autonomous DrivingabstractExplainability is essential for autonomous vehicles and other robotics systems interacting with humans and other objects during operation. Humans need to understand and anticipate the actions taken by machines for trustful and safe cooperation. In this work, we aim to develop an explainable model that generates explanations consistent with both human domain knowledge and the model’s inherent causal relation. In particular, we focus on an essential building block of autonomous driving—multi-agent interaction modeling. We propose Grounded Relational Inference (GRI). It models an interactive system’s underlying dynamics by inferring an interaction graph representing the agents’ relations. We ensure a semantically meaningful interaction graph by grounding the relational latent space into semantic interactive behaviors defined with expert domain knowledge. We demonstrate that it can model interactive traffic scenarios under both simulation and real-world settings and generate semantic graphs explaining the vehicle’s behavior through their interactions. Chen Tang 0001, Nishan Srishankar, Sujitha Martin, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Clothoid-Based Reference Path Reconstruction for HD Map GenerationabstractHigh-definition (HD) map is one of the key assets for autonomous driving, which supports various modules such as behavior prediction and motion planning of autonomous vehicles by providing accurate and rich geometric and semantic information. However, at present, the scalability and computational efficiency of HD map generation cannot meet the needs of highly automated driving. Specifically, efficiently obtaining the optimal parameters of the road’s reference path is still an open problem. In this paper, we propose a fast and robust path reconstruction method, which compresses the dense points of a reference line into sparse parameters with minimal loss of information. The reconstructed path consists of segmented linear curvature contours, which are straight lines, circular arcs, and clothoids. The optimum result is obtained through linear programming for short-path reconstruction, and for the long paths, a fast progressive reconstruction approach is used to find a feasible solution. Experimental results on both randomly generated data and the GPS-collected trajectories show that compared with existing methods, the proposed method can generate more accurate path reconstruction, and the computational time is greatly reduced. Songyi Zhang, Runsheng Wang, Zhiqiang Jian, Nanning Zheng 0001, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | RoadBEV: Road Surface Reconstruction in Bird's Eye ViewabstractRoad surface conditions, especially geometry profiles, enormously affect driving performance of autonomous vehicles. Vision-based online road reconstruction promisingly captures road information in advance. Existing solutions like monocular depth estimation and stereo matching suffer from modest performance. The recent technique of Bird’s-Eye-View (BEV) perception provides immense potential to more reliable and accurate reconstruction. This paper uniformly proposes two simple yet effective models for road elevation reconstruction in BEV named RoadBEV-mono and RoadBEV-stereo, which estimate road elevation with monocular and stereo images, respectively. The former directly fits elevation values based on voxel features queried from image view, while the latter efficiently recognizes road elevation patterns based on BEV volume representing correlation between left and right voxel features. Insightful analyses reveal their consistence and difference with the perspective view. Experiments on real-world dataset verify the models’ effectiveness and superiority. Elevation errors of RoadBEV-mono and RoadBEV-stereo achieve 1.83 cm and 0.50 cm, respectively. Our models are promising for practical road preview, providing essential information for promoting safety and comfort of autonomous vehicles. The code is released athttps://github.com/ztsrxh/RoadBEV. Lei Yang 0060, Yichen Xie 0002, Mingyu Ding, Masayoshi Tomizuka, Yintao Wei |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Design, Control, and Validation of a Novel Cable-Driven Series Elastic Actuation System for a Flexible and Portable Back-Support ExoskeletonabstractVarious active back-support exoskeletons have been developed to assist manual materials handling work for low back injury prevention. Existing back-support exoskeleton actuation either suffers from rigid transmission structure, or fails to efficiently generate assistance via portable actuation system with flexible transmissions. In this paper, a novel cable-driven series elastic actuation (CSEA) system is proposed to realize a flexible and portable back-support exoskeleton design with safe, efficient, and sufficient assistive torque output capability. The CSEA system realizes a flexible actuation based on cable transmission for an ergonomic human-exoskeleton interaction. Based on a torsion spring-support beam mechanism, it achieves an efficient assistance output capability to prevent high cable force demand and resultant lumbar compression, assuring a safe and synergistic operation for flexible exoskeleton actuation. Meanwhile, this mechanism enables the CSEA system to integrate series elastic actuator (SEA) with cable transmission and operates with multiple statuses to leverage SEA advantages and to overcome its torque output limitation. Dynamic model is established for the CSEA system, and a unified torque controller is designed for stable, continuous, and accurate torque control of the CSEA system despite its discontinuous dynamics during operation status transition. The efficacy of the closed-loop CSEA system to enable an ergonomic and efficient back-support exoskeleton actuation with the capability of accurately delivering desired level of assistance is verified via bench tests and human tests. Results verified that the CSEA system actuated exoskeleton can effectively reduce activity of relevant muscles during trunk flexion and extension motions compared to no exoskeleton case, validating successful application of the CSEA system on the exoskeleton for an effective back support effect. Hongpeng Liao, Hugo Hung-tin Chan, Gaoyu Liu, Xuan Zhao 0019, Masayoshi Tomizuka, Wei-Hsin Liao |
IEEE Trans. Robotics | 6 |
| 2023 | Open-Vocabulary Point-Cloud Object Detection without 3D AnnotationabstractThe goal of open-vocabulary detection is to identify novel objects based on arbitrary textual descriptions. In this paper, we address open-vocabulary 3D point-cloud detection by a dividing-and-conquering strategy, which involves: 1) developing a point-cloud detector that can learn a general representation for localizing various objects, and 2) connecting textual and point-cloud representations to enable the detector to classify novel object categories based on text prompting. Specifically, we resort to rich image pretrained models, by which the point-cloud detector learns localizing objects under the supervision of predicted 2D bounding boxes from 2D pretrained detectors. Moreover, we propose a novel de-biased triplet cross-modal contrastive learning to connect the modalities of image, point-cloud and text, thereby enabling the point-cloud detector to benefit from vision-language pretrained models, i.e., CLIP. The novel use of image and vision-language pretrained models for point-cloud detectors allows for open-vocabulary 3D object detection without the need for 3D annotations. Experiments demonstrate that the proposed method improves at least 3.03 points and 7.47 points over a wide range of baselines on the ScanNet and SUN RGB-D datasets, respectively. Furthermore, we provide a comprehensive analysis to explain why our approach works. Code is available at https://github.com/lyhdet/OV-3DET Yuheng Lu, Chenfeng Xu, Xiaobao Wei, Masayoshi Tomizuka, Kurt Keutzer, Shanghang Zhang |
CVPR | 5 |
| 2023 | Active Finetuning: Exploiting Annotation Budget in the Pretraining-Finetuning ParadigmabstractGiven the large-scale data and the high annotation cost, pretraining-finetuning becomes a popular paradigm in multiple computer vision tasks. Previous research has covered both the unsupervised pretraining and supervised finetuning in this paradigm, while little attention is paid to exploiting the annotation budget for finetuning. To fill in this gap, we formally define this new active finetuning task focusing on the selection of samples for annotation in the pretraining-finetuning paradigm. We propose a novel method called ActiveFT for active finetuning task to select a subset of data distributing similarly with the entire unlabeled pool and maintaining enough diversity by optimizing a parametric model in the continuous space. We prove that the Earth Mover's distance between the distributions of the selected subset and the entire data pool is also reduced in this process. Extensive experiments show the leading performance and high efficiency of ActiveFT superior to baselines on both image classification and semantic segmentation. Our code is released at https://github.com/yichen928/ActiveFT. Yichen Xie 0002, Han Lu 0004, Junchi Yan, Xiaokang Yang 0001, Masayoshi Tomizuka |
CVPR | 5 |
| 2023 | Towards Modeling and Influencing the Dynamics of Human LearningabstractHumans have internal models of robots (like their physical capabilities), the world (like what will happen next), and their tasks (like a preferred goal). However, human internal models are not always perfect: for example, it is easy to underestimate a robot's inertia. Nevertheless, these models change and improve over time as humans gather more experience. Interestingly, robot actions influence what this experience is, and therefore influence how people's internal models change. In this work we take a step towards enabling robots to understand the influence they have, leverage it to better assist people, and help human models more quickly align with reality. Our key idea is to model the human's learning as a nonlinear dynamical system which evolves the human's internal model given new observations. We formulate a novel optimization problem to infer the human's learning dynamics from demonstrations that naturally exhibit human learning. We then formalize how robots can influence human learning by embedding the human's learning dynamics model into the robot planning problem. Although our formulations provide concrete problem statements, they are intractable to solve in full generality. We contribute an approximation that sacrifices the complexity of the human internal models we can represent, but enables robots to learn the nonlinear dynamics of these internal models. We evaluate our inference and planning methods in a suite of simulated environments and an in-person user study, where a 7DOF robotic arm teaches participants to be better teleoperators. While influencing human learning remains an open problem, our results demonstrate that this influence is possible and can be helpful in real human-robot interaction. Masayoshi Tomizuka, Anca D. Dragan, Andrea Bajcsy |
HRI | 2 |
| 2023 | DELFlow: Dense Efficient Learning of Scene Flow for Large-Scale Point CloudsabstractPoint clouds are naturally sparse, while image pixels are dense. The inconsistency limits feature fusion from both modalities for point-wise scene flow estimation. Previous methods rarely predict scene flow from the entire point clouds of the scene with one-time inference due to the memory inefficiency and heavy overhead from distance calculation and sorting involved in commonly used farthest point sampling, KNN, and ball query algorithms for local feature aggregation. To mitigate these issues in scene flow learning, we regularize raw points to a dense format by storing 3D coordinates in 2D grids. Unlike the sampling operation commonly used in existing works, the dense 2D representation 1) preserves most points in the given scene, 2) brings in a significant boost of efficiency, and 3) eliminates the density gap between points and pixels, allowing us to perform effective feature fusion. We also present a novel warping projection technique to alleviate the information loss problem resulting from the fact that multiple points could be mapped into one grid during projection when computing cost volume. Sufficient experiments demonstrate the efficiency and effectiveness of our method, outperforming the prior-arts on the FlyingThings3D and KITTI dataset. Our source codes will be released on https://github.com/IRMVLab/DELFlow. Chensheng Peng, Guangming Wang 0001, Xian Wan Lo, Xinrui Wu, Chenfeng Xu, Masayoshi Tomizuka, Hesheng Wang 0001 |
ICCV | 6 |
| 2023 | SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object DetectionabstractBy identifying four important components of existing LiDAR-camera 3D object detection methods (LiDAR and camera candidates, transformation, and fusion outputs), we observe that all existing methods either find dense candidates or yield dense representations of scenes. However, given that objects occupy only a small part of a scene, finding dense candidates and generating dense representations is noisy and inefficient. We propose SparseFusion, a novel multi-sensor 3D detection method that exclusively uses sparse candidates and sparse representations. Specifically, SparseFusion utilizes the outputs of parallel detectors in the LiDAR and camera modalities as sparse candidates for fusion. We transform the camera candidates into the LiDAR coordinate space by disentangling the object representations. Then, we can fuse the multi-modality candidates in a unified 3D space by a lightweight self-attention module. To mitigate negative transfer between modalities, we propose novel semantic and geometric cross-modality transfer modules that are applied prior to the modality-specific detectors. SparseFusion achieves state-of-the-art performance on the nuScenes benchmark while also running at the fastest speed, even outperforming methods with stronger backbones. We perform extensive experiments to demonstrate the effectiveness and efficiency of our modules and overall method pipeline. Our code will be made publicly available at https://github.com/yichen928/SparseFusion. Yichen Xie 0002, Chenfeng Xu, Marie-Julie Rakotosaona, Patrick Rim, Federico Tombari, Kurt Keutzer, Masayoshi Tomizuka |
ICCV | 7 |
| 2023 | NeRF-Det: Learning Geometry-Aware Volumetric Representation for Multi-View 3D Object DetectionabstractWe present NeRF-Det, a novel method for indoor 3D detection with posed RGB images as input. Unlike existing indoor 3D detection methods that struggle to model scene geometry, our method makes novel use of NeRF in an end-to-end manner to explicitly estimate 3D geometry, thereby improving 3D detection performance. Specifically, to avoid the significant extra latency associated with per-scene optimization of NeRF, we introduce sufficient geometry priors to enhance the generalizability of NeRF-MLP. Furthermore, we subtly connect the detection and NeRF branches through a shared MLP, enabling an efficient adaptation of NeRF to detection and yielding geometry-aware volumetric representations for 3D detection. Our method outperforms state-of-the-arts by 3.9 mAP and 3.1 mAP on the ScanNet and ARKITScenes benchmarks, respectively. We provide extensive analysis to shed light on how NeRF-Det works. As a result of our joint-training design, NeRF-Det is able to generalize well to unseen scenes for object detection, view synthesis, and depth estimation tasks without requiring per-scene optimization. Code is available at https://github.com/facebookresearch/NeRF-Det. Chenfeng Xu, Bichen Wu, Ji Hou, Sam S. Tsai, Ruilong Li, Jialiang Wang 0001, Peter Vajda, Kurt Keutzer, Masayoshi Tomizuka |
ICCV | 11 |
| 2023 | Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object Detection
Jinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer, Kris Makoto Kitani, Masayoshi Tomizuka |
ICLR | 6 |
| 2023 | AdaptDiffuser: Diffusion Models as Adaptive Self-evolving PlannersabstractDiffusion models have demonstrated their powerful generative capability in many tasks, with great potential to serve as a paradigm for offline reinforcement learning. However, the quality of the diffusion model is limited by the insufficient diversity of training data, which hinders the performance of planning and the generalizability to new tasks. This paper introduces AdaptDiffuser, an evolutionary planning method with diffusion that can self-evolve to improve the diffusion model hence a better planner, not only for seen tasks but can also adapt to unseen tasks. AdaptDiffuser enables the generation of rich synthetic expert data for goal-conditioned tasks using guidance from reward gradients. It then selects high-quality data via a discriminator to finetune the diffusion model, which improves the generalization ability to unseen tasks. Empirical experiments on two benchmark environments and two carefully designed unseen tasks in KUKA industrial robot arm and Maze2D environments demonstrate the effectiveness of AdaptDiffuser. For example, AdaptDiffuser not only outperforms the previous art Diffuser by 20.8% on Maze2D and 7.5% on MuJoCo locomotion, but also adapts better to new tasks, e.g., KUKA pick-and-place, by 27.9% without requiring additional expert data. More visualization results and demo videos could be found on our project page. Zhixuan Liang, Yao Mu 0001, Mingyu Ding, Fei Ni 0001, Masayoshi Tomizuka, Ping Luo 0002 |
ICML | 5 |
| 2023 | Center Feature Fusion: Selective Multi-Sensor Fusion of Center-based ObjectsabstractLeveraging multi-modal fusion, especially between camera and LiDAR, has become essential for building accurate and robust 3D object detection systems for autonomous vehicles. Until recently, point decorating approaches, in which point clouds are augmented with camera features, have been the dominant approach in the field. However, these approaches fail to utilize the higher resolution images from cameras. Recent works projecting camera features to the bird's-eye-view (BEV) space for fusion have also been proposed, however they require projecting millions of pixels, most of which only contain background information. In this work, we propose a novel approach Center Feature Fusion (CFF), in which we leverage center-based detection networks in both the camera and LiDAR streams to identify relevant object locations. We then use the center-based detection to identify the locations of pixel features relevant to object locations, a small fraction of the total number in the image. These are then projected and fused in the BEV frame. On the nuScenes dataset, we outperform the LiDAR-only baseline by 4.9% mAP while fusing up to 100x fewer features than other fusion methods. Philip L. Jacobson, Yiyang Zhou, Masayoshi Tomizuka, Ming C. Wu |
ICRA | 4 |
| 2023 | Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement LearningabstractHumans are capable of abstracting various tasks as different combinations of multiple attributes. This perspective of compositionality is vital for human rapid learning and adaption since previous experiences from related tasks can be combined to generalize across novel compositional settings. In this work, we aim to achieve zero-shot policy generalization of Reinforcement Learning (RL) agents by leveraging the task compositionality. Our proposed method is a meta-RL algorithm with disentangled task representation, explicitly encoding different aspects of the tasks. Policy generalization is then performed by inferring unseen compositional task representations via the obtained disentanglement without extra exploration. The evaluation is conducted on three simulated tasks and a challenging real-world robotic insertion task. Experimental results demonstrate that our proposed method achieves policy generalization to unseen compositional tasks in a zero-shot manner. Zheng Wu 0002, Yichen Xie 0002, Wenzhao Lian, Yanjiang Guo, Jianyu Chen 0002, Stefan Schaal, Masayoshi Tomizuka |
ICRA | 8 |
| 2023 | A Coarse-to-Fine Framework for Dual-Arm Manipulation of Deformable Linear Objects with Whole-Body Obstacle AvoidanceabstractManipulating deformable linear objects (DLOs) to achieve desired shapes in constrained environments with obstacles is a meaningful but challenging task. Global planning is necessary for such a highly-constrained task; however, accurate models of DLOs required by planners are difficult to obtain owing to their deformable nature, and the inevitable modeling errors significantly affect the planning results, probably resulting in task failure if the robot simply executes the planned path in an open-loop manner. In this paper, we propose a coarse-to-fine framework to combine global planning and local control for dual-arm manipulation of DLOs, capable of precisely achieving desired configurations and avoiding potential collisions between the DLO, robot, and obstacles. Specifically, the global planner refers to a simple yet effective DLO energy model and computes a coarse path to find a feasible solution efficiently; then the local controller follows that path as guidance and further shapes it with closed-loop feedback to compensate for the planning errors and improve the task accuracy. Both simulations and real-world experiments demonstrate that our framework can robustly achieve desired DLO configurations in constrained environments with imprecise DLO models, which may not be reliably achieved by only planning or control. Mingrui Yu 0001, Kangchen Lv, Masayoshi Tomizuka, Xiang Li 0009 |
ICRA | 4 |
| 2023 | Learning Generalizable Pivoting SkillsabstractThe skill of pivoting an object with a robotic system is challenging for the external forces that act on the system, mainly given by contact interaction. The complexity increases when the same skills are required to generalize across different objects. This paper proposes a framework for learning robust and generalizable pivoting skills, which consists of three steps. First, we learn a pivoting policy on an “unitary” object using Reinforcement Learning (RL). Then, we obtain the object's feature space by supervised learning to encode the kinematic properties of arbitrary objects. Finally, to adapt the unitary policy to multiple objects, we learn data-driven projections based on the object features to adjust the state and action space of the new pivoting task. The proposed approach is entirely trained in simulation. It requires only one depth image of the object and can zero-shot transfer to real-world objects. We demonstrate robustness to sim-to-real transfer and generalization to multiple objects. Xiang Zhang 0020, Siddarth Jain, Baichuan Huang, Masayoshi Tomizuka, Diego Romeres |
ICRA | 4 |
| 2023 | Allowing Safe Contact in Robotic Goal-Reaching: Planning and Tracking in Operational and Null SpacesabstractIn recent years, impressive results have been achieved in robotic manipulation. While many efforts focus on generating collision-free reference signals, few allow safe contact between the robot bodies and the environment. However, in human's daily manipulation, contact between arms and obstacles is prevalent and even necessary. This paper investigates the benefit of allowing safe contact during robotic manipulation and advocates generating and tracking compliance reference signals in both operational and null spaces. In addition, to optimize the collision-allowed trajectories, we present a hybrid solver that integrates sampling- and gradient-based approaches. We evaluate the proposed method on a goal-reaching task in five simulated and real-world environments with different collisional conditions. We show that allowing safe contact improves goal-reaching efficiency and provides feasible solutions in highly collisional scenarios where collision-free constraints cannot be enforced. Moreover, we demonstrate that planning in null space, in addition to operational space, improves trajectory safety. Further information is available at https://rolandzhu.github.io/ContactReach/. Xinghao Zhu, Wenzhao Lian, Bodi Yuan, C. Daniel Freeman, Masayoshi Tomizuka |
ICRA | 5 |
| 2023 | Residual Q-Learning: Offline and Online Policy Customization without ValueabstractImitation Learning (IL) is a widely used framework for learning imitative behavior from demonstrations. It is especially appealing for solving complex real-world tasks where handcrafting reward function is difficult, or when the goal is to mimic human expert behavior. However, the learned imitative policy can only follow the behavior in the demonstration. When applying the imitative policy, we may need to customize the policy behavior to meet different requirements coming from diverse downstream tasks. Meanwhile, we still want the customized policy to maintain its imitative nature. To this end, we formulate a new problem setting called policy customization. It defines the learning task as training a policy that inherits the characteristics of the prior policy while satisfying some additional requirements imposed by a target downstream task. We propose a novel and principled approach to interpret and determine the trade-off between the two task objectives. Specifically, we formulate the customization problem as a Markov Decision Process (MDP) with a reward function that combines 1) the inherent reward of the demonstration; and 2) the add-on reward specified by the downstream task. We propose a novel framework, Residual Q-learning, which can solve the formulated MDP by leveraging the prior policy without knowing the inherent reward or value function of the prior policy. We derive a family of residual Q-learning algorithms that can realize offline and online policy customization, and show that the proposed algorithms can effectively accomplish policy customization tasks in various environments. Demo videos and code are available on our website: https://sites.google.com/view/residualq-learning. Chenran Li, Chen Tang 0001, Haruki Nishimura, Jean Mercat, Masayoshi Tomizuka |
NeurIPS | 5 |
| 2023 | Towards Free Data Selection with General-Purpose ModelsabstractA desirable data selection algorithm can efficiently choose the most informative samples to maximize the utility of limited annotation budgets. However, current approaches, represented by active learning methods, typically follow a cumbersome pipeline that iterates the time-consuming model training and batch data selection repeatedly. In this paper, we challenge this status quo by designing a distinct data selection pipeline that utilizes existing general-purpose models to select data from various datasets with a single-pass inference without the need for additional training or supervision. A novel free data selection (FreeSel) method is proposed following this new pipeline. Specifically, we define semantic patterns extracted from inter-mediate features of the general-purpose model to capture subtle local information in each image. We then enable the selection of all data samples in a single pass through distance-based sampling at the fine-grained semantic pattern level. FreeSel bypasses the heavy batch selection process, achieving a significant improvement in efficiency and being 530x faster than existing active learning methods. Extensive experiments verify the effectiveness of FreeSel on various computer vision tasks. Yichen Xie 0002, Mingyu Ding, Masayoshi Tomizuka |
NeurIPS | 3 |
| 2023 | Sparse R-CNN: An End-to-End Framework for Object DetectionabstractObject detection serves as one of most fundamental computer vision tasks. Existing works on object detection heavily rely on dense object candidates, such as k anchor boxes pre-defined on all grids of an image feature map of size H×W. In this paper, we present Sparse R-CNN, a very simple and sparse method for object detection in images. In our method, a fixed sparse set of learned object proposals ( N in total) are provided to the object recognition head to perform classification and localization. By replacing HWk (up to hundreds of thousands) hand-designed object candidates with N (e.g., 100) learnable proposals, Sparse R-CNN makes all efforts related to object candidates design and one-to-many label assignment completely obsolete. More importantly, Sparse R-CNN directly outputs predictions without the non-maximum suppression (NMS) post-processing procedure. Thus, it establishes an end-to-end object detection framework. Sparse R-CNN demonstrates highly competitive accuracy, run-time and training convergence performance with the well-established detector baselines on the challenging COCO dataset and CrowdHuman dataset. We hope that our work can inspire re-thinking the convention of dense prior in object detectors and designing new high-performance detectors. Peize Sun, Rufeng Zhang, Yi Jiang 0009, Tao Kong, Chenfeng Xu, Masayoshi Tomizuka, Zehuan Yuan, Ping Luo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | On Symmetric Gauss-Seidel ADMM Algorithm for H∞ Guaranteed Cost Control With Convex ParameterizationabstractThis article involves the innovative development of a symmetric Gauss–Seidel ADMM algorithm to solve the$\mathcal {H}_{\infty }$guaranteed cost control problem. In the presence of parametric uncertainties, the$\mathcal {H}_{\infty }$guaranteed cost control problem generally leads to the large-scale optimization. This is due to the exponential growth of the number of the extreme systems involved with respect to the number of parametric uncertainties. In this work, through a variant of the Youla–Kucera parameterization, the stabilizing controllers are parameterized in a convex set; yielding the outcome that the$\mathcal {H}_{\infty }$guaranteed cost control problem is converted to a convex optimization problem. Based on an appropriate reformulation using the Schur complement, it then renders possible the use of the ADMM algorithm with symmetric Gauss–Seidel backward and forward sweeps. Significantly, this approach alleviates the often-times prohibitively heavy computational burden typical in many$\mathcal H_{\infty }$optimization problems while exhibiting good convergence guarantees, which is particularly essential for the related large-scale optimization procedures involved. With this approach, the desired robust stability is ensured, and the disturbance attenuation is maintained at the minimum level in the presence of parametric uncertainties. Rather importantly too, with the attained effectiveness, the methodology thus evidently possesses extensive applicability in various important controller synthesis problems, such as decentralized control, sparse control, and output feedback control problems. Jun Ma 0008, Zilong Cheng, Masayoshi Tomizuka, Tong Heng Lee |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Multi-Objective Diverse Human Motion Prediction with Knowledge DistillationabstractObtaining accurate and diverse human motion prediction is essential to many industrial applications, especially robotics and autonomous driving. Recent research has ex-plored several techniques to enhance diversity and maintain the accuracy of human motion prediction at the same time. However, most of them need to define a combined loss, such as the weighted sum of accuracy loss and diversity loss, and then decide their weights as hyperparameters before training. In this work, we aim to design a prediction frame-work that can balance the accuracy sampling and diversity sampling during the testing phase. In order to achieve this target, we propose a multi-objective conditional variational inference prediction model. We also propose a short-term oracle to encourage the prediction framework to explore more diverse future motions. We evaluate the performance of our proposed approach on two standard human motion datasets. The experiment results show that our approach is effective and on a par with state-of-the-art performance in terms of accuracy and diversity. Hengbo Ma, Jiachen Li 0001, Ramtin Hosseini, Masayoshi Tomizuka, Chiho Choi |
CVPR | 4 |
| 2022 | What Matters for 3D Scene Flow Network
Guangming Wang 0001, Yunzhe Hu, Zhe Liu 0022, Yiyang Zhou, Masayoshi Tomizuka, Hesheng Wang 0001 |
ECCV (33) | 5 |
| 2022 | DetMatch: Two Teachers are Better than One for Joint 2D and 3D Semi-Supervised Object Detection
Jinhyung Park, Chenfeng Xu, Yiyang Zhou, Masayoshi Tomizuka |
ECCV (10) | 4 |
| 2022 | PreTraM: Self-supervised Pre-training via Connecting Trajectory and Map
Chenfeng Xu, Chen Tang 0001, Lingfeng Sun, Kurt Keutzer, Masayoshi Tomizuka, Alireza Fathi |
ECCV (39) | 6 |
| 2022 | Image2Point: 3D Point-Cloud Understanding with 2D Image Pretrained Models
Chenfeng Xu, Shijia Yang, Tomer Galanti, Bichen Wu, Xiangyu Yue 0001, Bohan Zhai, Peter Vajda, Kurt Keutzer, Masayoshi Tomizuka |
ECCV (37) | 10 |
| 2022 | Important Object Identification with Semi-Supervised Learning for Autonomous DrivingabstractAccurate identification of important objects in the scene is a prerequisite for safe and high-quality decision making and motion planning of intelligent agents (e.g., autonomous vehicles) that navigate in complex and dynamic environments. Most existing approaches attempt to employ attention mechanisms to learn importance weights associated with each object indirectly via various tasks (e.g., trajectory prediction), which do not enforce direct supervision on the importance estimation. In contrast, we tackle this task in an explicit way and formulate it as a binary classification (“important” or “unimportant”) problem. We propose a novel approach for important object identification in egocentric driving scenarios with relational reasoning on the objects in the scene. Besides, since human annotations are limited and expensive to obtain, we present a semi-supervised learning pipeline to enable the model to learn from unlimited unlabeled data. Moreover, we propose to leverage the auxiliary tasks of ego vehicle behavior prediction to further improve the accuracy of importance estimation. The proposed approach is evaluated on a public egocentric driving dataset (H3D) collected in complex traffic scenarios. A detailed ablative study is conducted to demonstrate the effectiveness of each model component and the training strategy. Our approach also outperforms rule-based baselines by a large margin. Jiachen Li 0001, Haiming Gang, Hengbo Ma, Masayoshi Tomizuka, Chiho Choi |
ICRA | 4 |
| 2022 | Causal-based Time Series Domain Generalization for Vehicle Intention PredictionabstractAccurately predicting the possible behaviors of traffic participants is an essential capability for autonomous vehicles. Since autonomous vehicles need to navigate in dynamically changing environments, they are expected to make accurate predictions regardless of where they are and what driving circumstances they encountered. Therefore, generalization capability to unseen domains is crucial for prediction models when autonomous vehicles are deployed in the real world. In this paper, we aim to address the domain generalization problem for vehicle intention prediction tasks and a causal-based time series domain generalization (CTSDG) model is proposed. We construct a structural causal model for vehicle intention prediction tasks to learn an invariant representation of input driving data for domain generalization. We further integrate a recurrent latent variable model into our structural causal model to better capture temporal latent dependencies from time-series input data. The effectiveness of our approach is evaluated via real-world driving data. We demonstrate that our proposed method has consistent improvement on prediction accuracy compared to other state-of-the-art domain generalization and behavior prediction methods. Yeping Hu, Xiaogang Jia, Masayoshi Tomizuka |
ICRA | 3 |
| 2022 | Autonomous Vehicle Parking in Dynamic Environments: An Integrated System with Prediction and Motion PlanningabstractThis paper presents an integrated motion planning system for autonomous vehicle (AV) parking in the presence of other moving vehicles. The proposed system includes 1) a hybrid environment predictor that predicts the motions of the surrounding vehicles and 2) a strategic motion planner that reacts to the predictions. The hybrid environment predictor performs short-term predictions via an extended Kalman filter and an adaptive observer. It also combines short-term predictions with a driver behavior cost-map to make long-term predictions. The strategic motion planner comprises 1) a model predictive control-based safety controller for trajectory tracking; 2) a search-based retreating planner for finding an evasion path in an emergency; 3) an optimization-based repairing planner for planning a new path when the original path is invalidated. Simulation validation demonstrates the effectiveness of the proposed method in terms of initial planning, motion prediction, safe tracking, retreating in an emergency, and trajectory repairing. Jessica EnShiuan Leu, Yebin Wang, Masayoshi Tomizuka, Stefano Di Cairano |
ICRA | 3 |
| 2022 | Cost-Effective Sensing for Goal Inference: A Model Predictive ApproachabstractGoal inference is of great importance for a variety of applications that involve interaction, coordination, and/or competition with goal-oriented agents. Typical goal inference approaches use as many pointwise measurements of the agent's trajectory as possible to pursue a most accurate a-posteriori estimate of the goal. However, taking frequent measurements may not be preferred in situations where sensing is associated with high cost (e.g., sensing + perception may involve high computational/bandwidth cost and sensing may raise security concerns in privacy-critical/data-sensitive applications). In such situations, a sensible tradeoff between the information gained from measurements and the cost associated with sensing actions is highly desirable. This paper introduces a cost-effective sensing strategy for goal inference tasks based on hybrid Kalman filtering and model predictive control. Our key insights include: 1) a model predictive approach can be used to predict the amount of information gained from new measurements over a horizon and thus to optimize the tradeoff between information gain and sensing action cost, and 2) the high computational efficiency of hybrid Kalman filtering can ensure real-time feasibility of such a model predictive approach. We evaluate the proposed cost-effective sensing approach in a goal-oriented task, where we show that compared to standard goal inference approaches, our approach takes a considerably reduced number of measurements while not impairing the speed, accuracy, and reliability of goal inference by taking measurements smartly. Nan Li 0015, Anouck R. Girard, Ilya V. Kolmanovsky, Masayoshi Tomizuka |
ICRA | 5 |
| 2022 | Safety Assurances for Human-Robot Interaction via Confidence-aware Game-theoretic Human ModelsabstractAn outstanding challenge with safety methods for human-robot interaction is reducing their conservatism while maintaining robustness to variations in human behavior. In this work, we propose that robots use confidence-aware game-theoretic models of human behavior when assessing the safety of a human-robot interaction. By treating the influence between the human and robot as well as the human's rationality as unobserved latent states, we succinctly infer the degree to which a human is following the game-theoretic interaction model. We leverage this model to restrict the set of feasible human controls during safety verification, enabling the robot to confidently modulate the conservatism of its safety monitor online. Evaluations in simulated human-robot scenarios and ablation studies demonstrate that imbuing safety monitors with confidence-aware game-theoretic models enables both safe and efficient human-robot interaction. Moreover, evaluations with real traffic data show that our safety monitor is less conservative than traditional safety methods in real human driving scenarios. Andrea Bajcsy, Masayoshi Tomizuka, Anca D. Dragan |
ICRA | 4 |
| 2022 | Cross Domain Robot Imitation with Invariant RepresentationabstractAnimals are able to imitate each others' behavior, despite their difference in biomechanics. In contrast, imitating other similar robots is a much more challenging task in robotics. This problem is called cross domain imitation learning (CDIL). In this paper, we consider CDIL on a class of similar robots. We tackle this problem by introducing an imitation learning algorithm based on invariant representation. We propose to learn invariant state and action representations, which align the behavior of multiple robots so that CDIL becomes possible. Compared with previous invariant representation learning methods for similar purposes, our method does not require human-labeled pairwise data for training. Instead, we use cycle-consistency and domain confusion to align the representation and increase its robustness. We test the algorithm on multiple robots in the simulator and show that unseen new robot instances can be trained with existing expert demonstrations successfully. Qualitative results also demonstrate that the proposed method is able to learn similar representations for different robots with similar behaviors, which is essential for successful CDIL. Zhao-Heng Yin, Lingfeng Sun, Hengbo Ma, Masayoshi Tomizuka, Wu-Jun Li |
ICRA | 4 |
| 2022 | Learning Insertion Primitives with Discrete-Continuous Hybrid Action Space for Robotic Assembly TasksabstractThis paper introduces a discrete-continuous action space to learn insertion primitives for robotic assembly tasks. Primitives are sequences of elementary actions with certain exit conditions, such as “pushing down the peg until contact”. Since the primitive is an abstraction of robot control commands and encodes human prior knowledge, it reduces the exploration difficulty and yields better learning efficiency. In this paper, we learn robot assembly skills via primitives. Specifically, we formulate insertion primitives as parameterized actions: hybrid actions consisting of discrete primitive types and continuous primitive parameters. Compared with the previous work using a set of discretized parameters for each primitive, the agent in our method can freely choose primitive parameters from a continuous space, which is more flexible and efficient. To learn these insertion primitives, we propose Twin-Smoothed Multi-pass Deep Q-Network (TS-MP-DQN), an advanced version of MP-DQN with twin Q-network to reduce the Q-value over-estimation. Extensive experiments are conducted in the simulation and real world for validation. From experiment results, our approach achieves higher success rates than three baselines: MP-DQN with parameterized actions, primitives with discrete parameters, and continuous velocity control. Furthermore, learned primitives are robust to sim-to-real transfer and can generalize to challenging assembly tasks such as tight round peg-hole and complex shaped electric connectors with promising success rates. Experiment videos are available at https://msc.berkeley.edu/research/insertion-primitives.html. Xiang Zhang 0020, Shiyu Jin, Xinghao Zhu, Masayoshi Tomizuka |
ICRA | 5 |
| 2022 | Grouptron: Dynamic Multi-Scale Graph Convolutional Networks for Group-Aware Dense Crowd Trajectory ForecastingabstractAccurate, long-term forecasting of pedestrian trajectories in highly dynamic and interactive scenes is a longstanding challenge. Recent advances in using data-driven approaches have achieved significant improvements in terms of prediction accuracy. However, the lack of group-aware analysis has limited the performance of forecasting models. This is especially nonnegligible in highly crowded scenes, where pedestrians are moving in groups and the interactions between groups are extremely complex and dynamic. In this paper, we present Grouptron, a multi-scale dynamic forecasting framework that leverages pedestrian group detection and utilizes individual-level, group-level and scene-level information for better understanding and representation of the scenes. Our approach employs spatio-temporal clustering algorithms to identify pedestrian groups, creates spatio-temporal graphs at the individual, group, and scene levels. It then uses graph neural networks to encode dynamics at different scales and aggregate the embeddings for trajectory prediction. We conducted extensive comparisons and ablation experiments to demonstrate the effectiveness of our approach. Our method achieves 9.3% decrease in final displacement error (FDE) compared with state-of-the-art methods on ETH/UCY benchmark datasets, and 16.1% decrease in FDE in more crowded scenes where extensive human group interactions are more frequently present. Huidong Gao, Masayoshi Tomizuka, Jiachen Li 0001 |
ICRA | 4 |
| 2022 | Learning to Synthesize Volumetric Meshes from Vision-based Tactile ImprintsabstractVision-based tactile sensors typically utilize a deformable elastomer and a camera mounted above to provide high-resolution image observations of contacts. Obtaining accurate volumetric meshes for the deformed elastomer can provide direct contact information and benefit robotic grasping and manipulation. This paper focuses on learning to synthesize the volumetric mesh of the elastomer based on the image imprints acquired from vision-based tactile sensors. Synthetic image-mesh pairs and real-world images are gathered from 3D finite element methods (FEM) and physical sensors, respectively. A graph neural network (GNN) is introduced to learn the image-to-mesh mappings with supervised learning. A self-supervised adaptation method and image augmentation techniques are proposed to transfer networks from simulation to reality, from primitive contacts to unseen contacts, and from one sensor to another. Using these learned and adapted networks, our proposed method can accurately reconstruct the deformation of the real-world tactile sensor elastomer in various domains, as indicated by the quantitative and qualitative results. Xinghao Zhu, Siddarth Jain, Masayoshi Tomizuka, Jeroen van Baar |
ICRA | 3 |
| 2022 | Learn to Grasp with Less Supervision: A Data-Efficient Maximum Likelihood Grasp Sampling LossabstractRobotic grasping for a diverse set of objects is essential in many robot manipulation tasks. One promising approach is to learn deep grasping models from large training datasets of object images and grasp labels. However, empirical grasping datasets are typically sparsely labeled (i.e., a small number of successful grasp labels**Labels refer to marking the image to indicate a successful robotic grasp. in each image). The data sparsity issue can lead to insufficient supervision and false-negative labels, and thus results in poor learning results. This paper proposes a Maximum Likelihood Grasp Sampling Loss (MLGSL) to tackle the data sparsity issue. The proposed method supposes that successful grasps are stochastically sampled from the predicted grasp distribution and maximizes the observing likelihood. MLGSL is utilized for training a fully convolutional network that generates thousands of grasps simultaneously. Training results suggest that models based on MLGSL can learn to grasp with datasets composing of 2 labels per image. Compared to previous works, which require training datasets of 16 labels per image, MLGSL is 8× more data-efficient. Meanwhile, physical robot experiments demonstrate an equivalent performance at a 90.7% grasp success rate on household objects. Codes and videos are available at [1]. Xinghao Zhu, Yefan Zhou, Yongxiang Fan, Lingfeng Sun, Jianyu Chen 0002, Masayoshi Tomizuka |
ICRA | 6 |
| 2022 | Improved A-Search Guided Tree for Autonomous Trailer PlanningabstractThis paper presents a motion planning strategy that utilizes the improved A -search guided tree to enable autonomous parking of a general 3-trailer with a car-like tractor. Different from the state-of-the-art state-lattice-based methods, where numerous motion primitives are necessary to ensure successful planning, our work allows quick off-lattice exploration to find a solution. Our treatment brings at least three advantages: fewer and lower design complexity of motion primitives, improved success rate, and increased path quality. Unlike on-lattice exploration, where the cost-to-go is obtained by querying a heuristic look-up table, off-lattice exploration entails the heuristic function being well-defined at off-lattice nodes. We train a neural network through reinforcement learning to model the maneuver costs of the trailer and use it as the heuristic value to better approximate the cost-to-go. Simulations demonstrate the effectiveness of the proposed method in terms of planning speed and path length. Jessica EnShiuan Leu, Yebin Wang, Masayoshi Tomizuka, Stefano Di Cairano |
IROS | 3 |
| 2022 | Generalizability Analysis of Graph-based Trajectory Predictor with Vectorized RepresentationabstractTrajectory prediction is one of the essential tasks for autonomous vehicles. Recent progress in machine learning gave birth to a series of advanced trajectory prediction algorithms. Lately, the effectiveness of using graph neural networks (GNNs) with vectorized representations for trajec-tory prediction has been demonstrated by many researchers. Nonetheless, these algorithms either pay little attention to models' generalizability across various scenarios or simply assume training and test data follow similar statistics. In fact, when test scenarios are unseen or Out-of-Distribution (OOD), the resulting train-test domain shift usually leads to significant degradation in prediction performance, which will impact downstream modules and eventually lead to severe accidents. Therefore, it is of great importance to thoroughly investigation of the prediction models in terms of their generalizability, which can not only help identify their weaknesses but also provide insights on how to improve these models. This paper proposes a generalizability analysis framework using feature attribution methods to help interpret black-box models. For the case study, we provide an in-depth generalizability analysis of one of the state-of-the-art graph-based trajectory predictors that utilize vectorized representation. Results show significant performance degradation due to domain shift, and feature attribution provides insights to identify potential causes of these problems. Finally, we conclude the common prediction challenges and how weighting biases induced by the training process can deteriorate the accuracy. Juanwu Lu, Masayoshi Tomizuka, Yeping Hu |
IROS | 3 |
| 2022 | Domain Knowledge Driven Pseudo Labels for Interpretable Goal-Conditioned Interactive Trajectory PredictionabstractMotion forecasting in highly interactive scenarios is a challenging problem in autonomous driving. In such scenarios, we need to accurately predict the joint behavior of interacting agents to ensure the safe and efficient navigation of autonomous vehicles. Recently, goal-conditioned methods have gained increasing attention due to their advantage in performance and their ability to capture the multimodality in trajec-tory distribution. In this work, we study the joint trajectory prediction problem with the goal-conditioned framework. In particular, we introduce a conditional-variational-autoencoder-based (CVAE) model to explicitly encode different interaction modes into the latent space. However, we discover that the vanilla model suffers from posterior collapse and cannot induce an informative latent space as desired. To address these issues, we propose a novel approach to avoid KL vanishing and induce an interpretable interactive latent space with pseudo labels. The proposed pseudo labels allow us to incorporate domain knowledge on interaction in a flexible manner. We motivate the proposed method using an illustrative toy example. In addition, we validate our framework on the Waymo Open Motion Dataset with both quantitative and qualitative evaluations. Lingfeng Sun, Chen Tang 0001, Yaru Niu, Enna Sachdeva, Chiho Choi, Teruhisa Misu, Masayoshi Tomizuka |
IROS | 7 |
| 2022 | Interventional Behavior Prediction: Avoiding Overly Confident Anticipation in Interactive PredictionabstractConditional behavior prediction (CBP) builds up the foundation for a coherent interactive prediction and plan-ning framework that can enable more efficient and less conser-vative maneuvers in interactive scenarios. In CBP task, we train a prediction model approximating the posterior distribution of target agents' future trajectories conditioned on the future trajectory of an assigned ego agent. However, we argue that CBP may provide overly confident anticipation on how the autonomous agent may influence the target agents' behavior. Consequently, it is risky for the planner to query a CBP model. Instead, we should treat the planned trajectory as an intervention and let the model learn the trajectory distribution under intervention. We refer to it as the interventional behavior prediction (IBP) task. Moreover, to properly evaluate an IBP model with offline datasets, we propose a Shapley-value-based metric to verify if the prediction model satisfies the inherent temporal independence of an interventional distribution. We show that the proposed metric can effectively identify a CBP model violating the temporal independence, which plays an important role when establishing IBP benchmarks. Chen Tang 0001, Masayoshi Tomizuka |
IROS | 3 |
| 2022 | PaCo: Parameter-Compositional Multi-task Reinforcement LearningabstractThe purpose of multi-task reinforcement learning (MTRL) is to train a single policy that can be applied to a set of different tasks. Sharing parameters allows us to take advantage of the similarities among tasks. However, the gaps between contents and difficulties of different tasks bring us challenges on both which tasks should share the parameters and what parameters should be shared, as well as the optimization challenges due to parameter sharing. In this work, we introduce a parameter-compositional approach (PaCo) as an attempt to address these challenges. In this framework, a policy subspace represented by a set of parameters is learned. Policies for all the single tasks lie in this subspace and can be composed by interpolating with the learned set. It allows not only flexible parameter sharing, but also a natural way to improve training.We demonstrate the state-of-the-art performance on Meta-World benchmarks, verifying the effectiveness of the proposed approach. Lingfeng Sun, Haichao Zhang 0001, Wei Xu 0017, Masayoshi Tomizuka |
NeurIPS | 4 |
| 2022 | AutoScale: Learning to Scale for Crowd Counting
Chenfeng Xu, Dingkang Liang, Yongchao Xu, Song Bai 0001, Xiang Bai, Masayoshi Tomizuka |
Int. J. Comput. Vis. | 7 |
| 2022 | Interpretable End-to-End Urban Autonomous Driving With Latent Deep Reinforcement LearningabstractUnlike popular modularized framework, end-to-end autonomous driving seeks to solve the perception, decision and control problems in an integrated way, which can be more adapting to new scenarios and easier to generalize at scale. However, existing end-to-end approaches are often lack of interpretability, and can only deal with simple driving tasks like lane keeping. In this article, we propose an interpretable deep reinforcement learning method for end-to-end autonomous driving, which is able to handle complex urban scenarios. A sequential latent environment model is introduced and learned jointly with the reinforcement learning process. With this latent model, a semantic birdeye mask can be generated, which is enforced to connect with certain intermediate properties in today’s modularized framework for the purpose of explaining the behaviors of learned policy. The latent space also significantly reduces the sample complexity of reinforcement learning. Comparison tests in a realistic driving simulator show that the performance of our method in urban scenarios with crowded surrounding vehicles dominates many baselines including DQN, DDPG, TD3 and SAC. Moreover, through masked outputs, the learned model is able to provide a better explanation of how the car reasons about the driving environment. Jianyu Chen 0002, Shengbo Eben Li, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Labels are Not Perfect: Inferring Spatial Uncertainty in Object DetectionabstractThe availability of many real-world driving datasets is a key reason behind the recent progress of object detection algorithms in autonomous driving. However, there exist ambiguity or even failures in object labels due to error-prone annotation process or sensor observation noise. Current public object detection datasets only provide deterministic object labels without considering their inherent uncertainty, as does the common training process or evaluation metrics for object detectors. As a result, an in-depth evaluation among different object detection methods remains challenging, and the training process of object detectors is sub-optimal, especially in probabilistic object detection. In this work, we infer the uncertainty in bounding box labels from LiDAR point clouds based on a generative model, and define a new representation of the probabilistic bounding box through a spatial uncertainty distribution. Comprehensive experiments show that the proposed model reflects complex environmental noises in LiDAR perception and the label quality. Furthermore, we propose Jaccard IoU (JIoU) as a new evaluation metric that extends IoU by incorporating label uncertainty. We conduct an in-depth comparison among several LiDAR-based object detectors using the JIoU metric. Finally, we incorporate the proposed label uncertainty in a loss function to train a probabilistic object detector and to improve its detection accuracy. We verify our proposed methods on two public datasets (KITTI, Waymo), as well as on simulation data. Code is released athttps://github.com/ZiningWang/Inferring-Spatial-Uncertainty-in-Object-Detection. Di Feng, Yiyang Zhou, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | Scenario-Transferable Semantic Graph Reasoning for Interaction-Aware Probabilistic PredictionabstractAccurately predicting the possible behaviors of traffic participants is an essential capability for autonomous vehicles. Since autonomous vehicles need to navigate in dynamically changing environments, they are expected to make accurate predictions regardless of where they are and what driving circumstances they encountered. Several methodologies have been proposed to solve prediction problems under different traffic situations. These works usually combine agent trajectories with either color-coded or vectorized high definition (HD) map as input representations and encode this information for behavior prediction tasks. However, not all the information is relevant in the scene for the forecasting and such irrelevant information may be even distracting to the forecasting in certain situations. Therefore, in this paper, we propose a novel generic representation for various driving environments by taking the advantage of semantics and domain knowledge. Using semantics enables situations to be modeled in a uniform way and applying domain knowledge filters out unrelated elements to target vehicle’s future behaviors. We then propose a general semantic behavior prediction framework to effectively utilize these representations by formulating them into spatial-temporal semantic graphs and reasoning internal relations among these graphs. We theoretically and empirically validate the proposed framework under highly interactive and complex scenarios, demonstrating that our method not only achieves state-of-the-art performance, but also processes desirable zero-shot transferability. Yeping Hu, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | SceGene: Bio-Inspired Traffic Scenario Generation for Autonomous Driving TestingabstractThe core value of simulation-based autonomy tests is to create densely extreme traffic scenarios to test the performance and robustness of the algorithms and systems. Test scenarios are usually designed or extracted manually from the real-world data, which is inefficient with a remarkable domain gap compared with testing in real scenarios. Therefore, it is crucial to automatically generate realistic and diverse dynamic traffic scenarios making autonomy tests efficient. Moreover, scenario generation is expected to be interpretable, controllable, and diversified, which can be hard to achieve simultaneously by methods based on rules or deep networks. In this paper, we propose a dynamic traffic scenario generation method called SceGene, inspired by genetic inheritance and mutation processes in biological intelligence. SceGene applies biological processes, such as crossover and mutation, to exchange and mutate the content of scenarios, and involves the natural selection process to control generation direction. SceGene has three main parts: 1) a new representation method for describing the traffic scenarios’ feature; 2) a new scenario generation algorithm based on crossover, mutation, and selection; and 3) an abnormal scenario information repair method based on the microscopic driving model. Evaluation on the public traffic scenario dataset shows that SceGene can ensure highly realistic and diversified scenario generation in an interpretable and controllable way, significantly improving the efficiency of the simulation-based autonomy tests. Ao Li 0006, Shi-tao Chen, Nanning Zheng 0001, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Spatio-Temporal Graph Dual-Attention Network for Multi-Agent Prediction and TrackingabstractAn effective understanding of the environment and accurate trajectory prediction of surrounding dynamic obstacles are indispensable for intelligent mobile systems (e.g., autonomous vehicles, social robots) to achieve safe and high-quality planning when they navigate in highly interactive and crowded scenarios. Due to the existence of frequent interactions and uncertainty in the scene evolution, it is desired for the prediction system to enable relational reasoning on different entities and provide a distribution of future trajectories for each agent. In this paper, we propose a generic generative neural system (called STG-DAT) for multi-agent trajectory prediction involving heterogeneous agents. The system takes a step forward to explicit interaction modeling by incorporating relational inductive biases with a dynamic graph representation and leverages both trajectory and scene context information. We also employ an efficient kinematic constraint layer applied to vehicle trajectory prediction. The constraint not only ensures physical feasibility but also enhances model performance. Moreover, the proposed prediction model can be easily adopted by multi-target tracking frameworks. The tracking accuracy proves to be improved by empirical results. The proposed system is evaluated on three public benchmark datasets for trajectory prediction, where the agents cover pedestrians, cyclists and on-road vehicles. The experimental results demonstrate that our model achieves better performance than various baseline approaches in terms of prediction and tracking accuracy. Jiachen Li 0001, Hengbo Ma, Zhihao Zhang 0001, Jinning Li 0002, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Alternating Direction Method of Multipliers for Constrained Iterative LQR in Autonomous DrivingabstractIn the context of autonomous driving, the iterative linear quadratic regulator (iLQR) is known to be an efficient approach to deal with the nonlinear vehicle model in motion planning problems. Particularly, the constrained iLQR algorithm has shown noteworthy advantageous outcomes of computation efficiency in achieving motion planning tasks under general constraints of different types. However, the constrained iLQR methodology requires a feasible trajectory at the first iteration as a prerequisite when the logarithmic barrier function is used. Also, the methodology leaves open the possibility for incorporation of fast, efficient, and effective optimization methods (i.e., fast-solvers) to further speed up the optimization process such that the requirements of real-time implementation can be successfully fulfilled. In this paper, a well-defined and commonly-encountered motion planning problem is formulated under nonlinear vehicle dynamics and various constraints, and the alternating direction method of multipliers (ADMM) is utilized to determine the optimal control actions leveraging the iLQR. With this development, the approach is able to circumvent the feasibility requirement of the trajectory at the first iteration. An illustrative example of motion planning for autonomous vehicles is then investigated with different driving scenarios taken into consideration, and a noteworthy achievement of high computation efficiency is attained with the proposed development. Comparing with the constrained iLQR algorithm based on the logarithmic barrier function, our proposed method reduces the average computation time by 31.93%, 38.52%, and 44.57% in the three scenarios; compared with the optimization solver IPOPT, our proposed method reduces the average computation time by 46.02%, 53.26%, and 88.43% in the three scenarios. As a result, real-time computation and implementation can be realized through our proposed framework, and thus it provides additional safety to the on-road driving tasks. Jun Ma 0008, Zilong Cheng, Masayoshi Tomizuka, Tong Heng Lee |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | From Human Driving to Automated Driving: What Do We Know About Drivers?abstractHumanlike automated driving (AD) strategies which are inspired by drivers’ cognition ways may show advantages in dealing with complicated scenarios. However, many humanlike AD strategies just mimic drivers’ behaviors or some specific characteristic. Learning algorithms are powerful technics to realize these strategies, but the architectures in learning-based strategies are too simple or with no detailed foundations. Therefore, we mean to summarize drivers’ cognition characteristics and design a comprehensive and well-founded architecture for humanlike AD solutions. We review the massive studies about drivers with human driving or AD and summarize the characteristics from three perspectives, cognition foundation, cognition process, and cognition strategies. As for cognition foundation, we propose a simple analogy to show the working mechanisms of biological neural networks; as for cognition foundation, the important role of previous experience is highlighted; as for cognition strategies, we discuss drivers’ cognition compensation strategies under the influences of environment, vehicle automation, and personal states systematically. After the above review of drivers’ characteristics, we classify the methods to model drivers. We find that models based on cognition processes can maintain more cognition details, and thus we design a driving-dedicated cognitive architecture. This architecture works by the cooperation of several modules including long-term memory, management module, and so on. It has solid theoretical and factual foundations and can reflect drivers’ cognition characteristics comprehensively. Finally, we discuss what needs to be done in the near future for us to improve humanlike AD solutions gradually. Shi-tao Chen, Jingyue Zheng, Masayoshi Tomizuka, Nanning Zheng 0001, Jianqiang Wang 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | On Robust Stability and Performance With a Fixed-Order Controller Design for Uncertain SystemsabstractTypically, it is desirable to design a control system that is not only robustly stable in the presence of parametric uncertainties but also guarantees an adequate level of system performance. However, most of the existing methods need to take all extreme models over an uncertain domain into consideration, which then results in costly computation. Also, since these approaches attempt rather unrealistically to guarantee the system performance over a full frequency range, a conservative design is always admitted. Here, taking a specific viewpoint of robust stability and performance under a stated restricted frequency range (which is applicable in rather many real-world situations), this article provides an essential basis for the design of a fixed-order controller for a system with bounded parametric uncertainties, which avoids the tedious but necessary evaluations of the specifications on all the extreme models in an explicit manner. A Hurwitz polynomial is used in the design and the robust stability is characterized by the notion of positive realness, such that the required robust stability condition is then successfully constructed. Also, the robust performance criteria in terms of sensitivity shaping under different frequency ranges are constructed based on an approach of bounded realness analysis. Furthermore, the conditions for robust stability and performance are expressed in the framework of linear matrix inequality (LMI) constraints, and thus can be efficiently solved. Comparative simulations are provided to demonstrate the effectiveness and efficiency of the proposed approach. Jun Ma 0008, Haiyue Zhu, Masayoshi Tomizuka, Tong Heng Lee |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Bounded Risk-Sensitive Markov Games: Forward Policy Design and Inverse Reward Learning with Iterative Reasoning and Cumulative Prospect TheoryabstractClassical game-theoretic approaches for multi-agent systems in both the forward policy design problem and the inverse reward learning problem often make strong rationality assumptions: agents perfectly maximize expected utilities under uncertainties. Such assumptions, however, substantially mismatch with observed human behaviors such as satisficing with sub-optimal, risk-seeking, and loss-aversion decisions. Drawing on iterative reasoning models and cumulative prospect theory, we propose a new game-theoretic framework, bounded risk-sensitive Markov Game (BRSMG), that captures two aspects of realistic human behaviors: bounded intelligence and risk-sensitivity. General solutions to both the forward policy design problem and the inverse reward learning problem are provided with theoretical analysis and simulation verification. We validate the proposed forward policy design algorithm and the inverse reward learning algorithm in a two-player navigation scenario. The results show that agents demonstrate bounded-intelligence, risk-averse and risk-seeking behaviors in our framework. Moreover, in the inverse reward learning task, the proposed bounded risk-sensitive inverse learning algorithm outperforms a baseline risk-neutral inverse learning algorithm by effectively learning not only more accurate reward values but also the intelligence levels and the risk-measure parameters of agents from demonstrations. Masayoshi Tomizuka |
AAAI | 3 |
| 2021 | Sparse R-CNN: End-to-End Object Detection With Learnable ProposalsabstractWe present Sparse R-CNN, a purely sparse method for object detection in images. Existing works on object detection heavily rely on dense object candidates, such as k anchor boxes pre-defined on all grids of image feature map of size H × W. In our method, however, a fixed sparse set of learned object proposals, total length of N, are provided to object recognition head to perform classification and location. By eliminating HWk (up to hundreds of thousands) hand-designed object candidates to N (e.g. 100) learnable proposals, Sparse R-CNN completely avoids all efforts related to object candidates design and many-to-one label assignment. More importantly, final predictions are directly output without non-maximum suppression post-procedure. Sparse R-CNN demonstrates accuracy, run-time and training convergence performance on par with the well-established detector baselines on the challenging COCO dataset, e.g., achieving 45.0 AP in standard 3× training schedule and running at 22 fps using ResNet-50 FPN model. We hope our work could inspire re-thinking the convention of dense prior in object detectors. The code is available at: https://github.com/PeizeSun/SparseR-CNN. Peize Sun, Rufeng Zhang, Yi Jiang 0009, Tao Kong, Chenfeng Xu, Masayoshi Tomizuka, Lei Li 0005, Zehuan Yuan, Changhu Wang, Ping Luo 0002 |
CVPR | 7 |
| 2021 | RAIN: Reinforced Hybrid Attention Inference Network for Motion ForecastingabstractMotion forecasting plays a significant role in various domains (e.g., autonomous driving, human-robot interaction), which aims to predict future motion sequences given a set of historical observations. However, the observed elements may be of different levels of importance. Some information may be irrelevant or even distracting to the forecasting in certain situations. To address this issue, we propose a generic motion forecasting framework (named RAIN) with dynamic key information selection and ranking based on a hybrid attention mechanism. The general framework is instantiated to handle multi-agent trajectory prediction and human motion forecasting tasks, respectively. In the former task, the model learns to recognize the relations between agents with a graph representation and to determine their relative significance. In the latter task, the model learns to capture the temporal proximity and dependency in long-term human motions. We also propose an effective double-stage training pipeline with an alternating training strategy to optimize the parameters in different modules of the framework. We validate the framework on both synthetic simulations and motion forecasting benchmarks in different domains, demonstrating that our method not only achieves state-of-the-art forecasting performance, but also provides interpretable and reasonable hybrid attention weights. Jiachen Li 0001, Hengbo Ma, Srikanth Malla, Masayoshi Tomizuka, Chiho Choi |
ICCV | 5 |
| 2021 | Visual Transformers: Where Do Transformers Really Belong in Vision Models?abstractA recent trend in computer vision is to replace convolutions with transformers. However, the performance gain of transformers is attained at a steep cost, requiring GPU years and hundreds of millions of samples for training. This excessive resource usage compensates for a misuse of transformers: Transformers densely model relationships between its inputs - ideal for late stages of a neural network, when concepts are sparse and spatially-distant, but extremely inefficient for early stages of a network, when patterns are redundant and localized. To address these issues, we leverage the respective strengths of both operations, building convolution-transformer hybrids. Critically, in sharp contrast to pixel-space transformers, our Visual Transformer (VT) operates in a semantic token space, judiciously attending to different image parts based on context. Our VTs significantly outperforms baselines: On ImageNet, our VT-ResNets outperform convolution-only ResNet by 4.6 to 7 points and transformer-only ViT-B by 2.6 points with 2.5× fewer FLOPs, 2.1× fewer parameters. For semantic segmentation on LIP and COCO-stuff, VT-based feature pyramid networks (FPN) achieve 0.35 points higher mIoU while reducing the FPN module’s FLOPs by 6.5x. Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan 0001, Masayoshi Tomizuka, Joseph Gonzalez 0001, Kurt Keutzer, Peter Vajda |
ICCV | 7 |
| 2021 | Spectral Temporal Graph Neural Network for Trajectory PredictionabstractAn effective understanding of the contextual environment and accurate motion forecasting of surrounding agents is crucial for the development of autonomous vehicles and social mobile robots. This task is challenging since the behavior of an autonomous agent is not only affected by its own intention, but also by the static environment and surrounding dynamically interacting agents. Previous works focused on utilizing the spatial and temporal information in time domain while not sufficiently taking advantage of the cues in frequency domain. To this end, we propose a Spectral Temporal Graph Neural Network (SpecTGNN), which can capture inter-agent correlations and temporal dependency simultaneously in frequency domain in addition to time domain. SpecTGNN operates on both an agent graph with dynamic state information and an environment graph with the features extracted from context images in two streams. The model integrates graph Fourier transform, spectral graph convolution and temporal gated convolution to encode history information and forecast future trajectories. Moreover, we incorporate a multi-head spatio-temporal attention mechanism to mitigate the effect of error propagation in a long time horizon. We demonstrate the performance of SpecTGNN on two public trajectory prediction benchmark datasets, which achieves state-of-the-art performance in terms of prediction accuracy. Defu Cao, Jiachen Li 0001, Hengbo Ma, Masayoshi Tomizuka |
ICRA | 4 |
| 2021 | Trajectory Optimization for Manipulation of Deformable Objects: Assembly of Belt Drive UnitsabstractThis paper presents a novel trajectory optimization formulation to solve the robotic assembly of the belt drive unit. Robotic manipulations involving contacts and deformable objects are challenging in both dynamic modeling and trajectory planning. For modeling, variations in the belt tension and contact forces between the belt and the pulley could dramatically change the system dynamics. For trajectory planning, it is computationally expensive to plan trajectories for such hybrid dynamical systems as it usually requires planning for discrete modes separately. In this work, we formulate the belt drive unit assembly task as a trajectory optimization problem with complementarity constraints to avoid explicitly imposing contact mode sequences. The problem is solved as a mathematical program with complementarity constraints (MPCC) to obtain feasible and efficient assembly trajectories. We validate the proposed method both in simulations with a physics engine and in real-world experiments with a robotic manipulator. Shiyu Jin, Diego Romeres, Arvind Ragunathan, Devesh K. Jha, Masayoshi Tomizuka |
ICRA | 5 |
| 2021 | A Safe Hierarchical Planning Framework for Complex Driving Scenarios based on Reinforcement LearningabstractAutonomous vehicles need to handle various traffic conditions and make safe and efficient decisions and maneuvers. However, on the one hand, a single optimization/sampling-based motion planner cannot efficiently generate safe trajectories in real time, particularly when there are many interactive vehicles near by. On the other hand, end-to-end learning methods cannot assure the safety of the outcomes. To address this challenge, we propose a hierarchical behavior planning framework with a set of low-level safe controllers and a high-level reinforcement learning algorithm (H-CtRL) as a coordinator for the low-level controllers. Safety is guaranteed by the low-level optimization/sampling-based controllers, while the high-level reinforcement learning algorithm makes H-CtRL an adaptive and efficient behavior planner. To train and test our proposed algorithm, we built a simulator that can reproduce traffic scenes using real-world datasets. The proposed HCtRL is proved to be effective in various realistic simulation scenarios, with satisfying performance in terms of both safety and efficiency. Jinning Li 0002, Jianyu Chen 0002, Masayoshi Tomizuka |
ICRA | 4 |
| 2021 | Prediction-Based Reachability for Collision Avoidance in Autonomous DrivingabstractSafety is an important topic in autonomous driving since any collision may cause serious injury to people and damage to property. Hamilton-Jacobi (HJ) Reachability is a formal method that verifies safety in multi-agent interaction and provides a safety controller for collision avoidance. However, due to the worst-case assumption on the car’s future behaviours, reachability might result in too much conservatism such that the normal operation of the vehicle is badly hindered. In this paper, we leverage the power of trajectory prediction and propose a prediction-based reachability framework to compute safety controllers. Instead of always assuming the worst case, we cluster the car’s behaviors into multiple driving modes, e.g. left turn or right turn. Under each mode, a reachability-based safety controller is designed based on a less conservative action set. For online implementation, we first utilize the trajectory prediction and our proposed mode classifier to predict the possible modes, and then deploy the corresponding safety controller. Through simulations in a T-intersection and an 8-way roundabout, we demonstrate that our prediction-based reachability method largely avoids collision between two interacting cars and reduces the conservatism that the safety controller brings to the car’s original operation. Anjian Li, Masayoshi Tomizuka, Mo Chen 0001 |
ICRA | 4 |
| 2021 | Anytime Game-Theoretic Planning with Active Reasoning About Humans' Latent States for Human-Centered RobotsabstractA human-centered robot needs to reason about the cognitive limitation and potential irrationality of its human partner to achieve seamless interactions. This paper proposes an anytime game-theoretic planner that integrates iterative reasoning models, a partially observable Markov decision process, and chance-constrained Monte-Carlo belief tree search for robot behavioral planning. Our planner enables a robot to safely and actively reason about its human partner’s latent cognitive states (bounded intelligence and irrationality) in real-time to maximize its utility better. We validate our approach in an autonomous driving domain where our behavioral planner and a low-level motion controller hierarchically control an autonomous car to negotiate traffic merges. Simulations and user studies are conducted to show our planner’s effectiveness. Masayoshi Tomizuka, David Isele |
ICRA | 3 |
| 2021 | Learning Dense Rewards for Contact-Rich Manipulation TasksabstractRewards play a crucial role in reinforcement learning. To arrive at the desired policy, the design of a suitable reward function often requires significant domain expertise as well as trial-and-error. Here, we aim to minimize the effort involved in designing reward functions for contact-rich manipulation tasks. In particular, we provide an approach capable of extracting dense reward functions algorithmically from robots’ high-dimensional observations, such as images and tactile feedback. In contrast to state-of-the-art high-dimensional reward learning methodologies, our approach does not leverage adversarial training, and is thus less prone to the associated training instabilities. Instead, our approach learns rewards by estimating task progress in a self-supervised manner. We demonstrate the effectiveness and efficiency of our approach on two contact-rich manipulation tasks, namely, peg-in-hole and USB insertion. The experimental results indicate that the policies trained with the learned reward function achieves better performance and faster convergence compared to the baselines. Zheng Wu 0002, Wenzhao Lian, Vaibhav V. Unhelkar, Masayoshi Tomizuka, Stefan Schaal |
ICRA | 4 |
| 2021 | 6-DoF Contrastive Grasp Proposal NetworkabstractProposing grasp poses for novel objects is an essential component for any robot manipulation task. Planning six degrees of freedom (DoF) grasps with a single camera, however, is challenging due to the complex object shape, incomplete object information, and sensor noise. In this paper, we present a 6-DoF contrastive grasp proposal network (CGPN) to infer 6-DoF grasps from a single-view depth image. First, an image encoder is used to extract the feature map from the input depth image, after which 3-DoF grasp regions are proposed from the feature map with a rotated region proposal network. Feature vectors that within the proposed grasp regions are then extracted and refined to 6-DoF grasps. The proposed model is trained offline with synthetic grasp data. To improve the robustness in reality and bridge the simulation-to-real gap, we further introduce a contrastive learning module and variant image processing techniques during the training. CGPN can locate collision-free grasps of an object using a single-view depth image within 0.5 second. Experiments on a physical robot further demonstrate the effectiveness of the algorithm. The experimental videos are available at [1]. Xinghao Zhu, Lingfeng Sun, Yongxiang Fan, Masayoshi Tomizuka |
ICRA | 4 |
| 2021 | Constrained Iterative LQG for Real-Time Chance-Constrained Gaussian Belief Space PlanningabstractMotion planning under uncertainty is of significant importance for safety-critical systems such as autonomous vehicles. Such systems have to satisfy necessary constraints (e.g., collision avoidance) with potential uncertainties coming from either disturbed system dynamics or noisy sensor measurements. However, existing motion planning methods cannot efficiently find the robust optimal solutions under general nonlinear and non-convex settings. In this paper, we formulate such problem as chance-constrained Gaussian belief space planning and propose the constrained iterative Linear Quadratic Gaussian (CILQG) algorithm as a real-time solution. In this algorithm, we iteratively calculate a Gaussian approximation of the belief and transform the chance-constraints. We evaluate the effectiveness of our method in simulations of autonomous driving planning tasks with static and dynamic obstacles. Results show that CILQG can handle uncertainties more appropriately and has faster computation time than baseline methods. Jianyu Chen 0002, Yutaka Shimizu, Masayoshi Tomizuka |
IROS | 4 |
| 2021 | A Simple and Efficient Multi-task Network for 3D Object Detection and Road UnderstandingabstractDetecting dynamic objects and predicting static road information such as drivable areas and ground heights are crucial for safe autonomous driving. Previous works studied each perception task separately, and lacked a collective quantitative analysis. In this work, we show that it is possible to perform all perception tasks via a simple and efficient multi-task network. Our proposed network, LidarMTL, takes raw LiDAR point cloud as inputs, and predicts six perception outputs for 3D object detection and road understanding. The network is based on an encoder-decoder architecture with 3D sparse convolution and deconvolution operations. Extensive experiments verify the proposed method with competitive accuracies compared to state-of-the-art object detectors and other task-specific networks. LidarMTL is also leveraged for online localization. Code and pre-trained model have been made available at https://github.com/frankfengdi/LidarMTL. Di Feng, Yiyang Zhou, Chenfeng Xu, Masayoshi Tomizuka |
IROS | 4 |
| 2021 | Learning Human Rewards by Inferring Their Latent Intelligence Levels in Multi-Agent Games: A Theory-of-Mind Approach with Application to Driving DataabstractReward function, as an incentive representation that recognizes humans’ agency and rationalizes humans’ actions, is particularly appealing for modeling human behavior in human-robot interaction. Inverse Reinforcement Learning is an effective way to retrieve reward functions from demonstrations. However, it has always been challenging when applying it to multi-agent settings since the mutual influence between agents has to be appropriately modeled. To tackle this challenge, previous work either exploits equilibrium solution concepts by assuming humans as perfectly rational optimizers with unbounded intelligence or pre-assigns humans’ interaction strategies a priori. In this work, we advocate that humans are bounded rational and have different intelligence levels when reasoning about others’ decision-making process, and such an inherent and latent characteristic should be accounted for in reward learning algorithms. Hence, we exploit such insights from Theory-of-Mind and propose a new multi-agent Inverse Reinforcement Learning framework that reasons about humans’ latent intelligence levels during learning. We validate our approach in both zero-sum and general-sum games with synthetic agents, and illustrate a practical application to learning human drivers’ reward functions from real driving data. We compare our approach with two baseline algorithms. The results show that by reasoning about humans’ latent intelligence levels, the proposed approach has more flexibility and capability to retrieve reward functions that explain humans’ driving behaviors better. Masayoshi Tomizuka |
IROS | 2 |
| 2021 | Trajectory Splitting: A Distributed Formulation for Collision Avoiding Trajectory OptimizationabstractEfficient trajectory optimization is essential for avoiding collisions in unstructured environments, but it remains challenging to have both speed and quality in the solutions. One reason is that second-order optimality requires calculating Hessian matrices that can grow with O(N2) with the number of waypoints. Decreasing the waypoints can quadratically decrease computation time. Unfortunately, fewer waypoints result in lower quality trajectories that may not avoid the collision. To have both, dense waypoints and reduced computation time, we took inspiration from recent studies on consensus optimization and propose a distributed formulation of collocated trajectory optimization. It breaks a long trajectory into several segments, where each segment becomes a subproblem of a few waypoints. These subproblems are solved classically, but in parallel, and the solutions are fused into a single trajectory with a consensus constraint that enforces continuity of the segments through a consensus update. With this scheme, the quadratic complexity is distributed to each segment and enables solving for higher-quality trajectories with denser waypoints. Furthermore, the proposed formulation is amenable to using any existing trajectory optimizer for solving the subproblems. We compare the performance of our implementation of trajectory splitting against leading motion planning algorithms and demonstrate the improved computational efficiency of our method. Jeffrey T. Bingham, Masayoshi Tomizuka |
IROS | 3 |
| 2021 | COCOI: Contact-aware Online Context Inference for Generalizable Non-planar PushingabstractGeneral contact-rich manipulation problems are long-standing challenges in robotics due to the difficulty of understanding complicated contact physics. Deep reinforcement learning (RL) has shown great potential in solving robot manipulation tasks. However, existing RL policies have limited adaptability to environments with diverse dynamics properties, which is pivotal in solving many contact-rich manipulation tasks. In this work, we propose Contact-aware Online COntext Inference (COCOI), a deep RL method that encodes a context embedding of dynamics properties online using contact-rich interactions. We sample sensor data using a novel contact-aware strategy and formulate an interpretable dynamics transition module. We study this method based on a novel and challenging non-planar pushing task, where the robot uses a monocular camera image and wrist force torque sensor reading to push an object to a goal location while keeping it upright. We run extensive experiments to demonstrate the capability of COCOI in a wide range of settings and dynamics properties in simulation, and also in a sim-to-real transfer scenario on a real robot (Webpage: https://context-inference.github.io/). Wenhao Yu 0003, Chuyuan Fu, Masayoshi Tomizuka, C. Karen Liu, Daniel Ho |
IROS | 6 |
| 2021 | You Only Group Once: Efficient Point-Cloud Processing with Token Representation and Relation Inference Moduleabstract3D perception on point-cloud is a challenging and crucial computer vision task. A point-cloud consists of a sparse, unstructured, and unordered set of points. To understand a point-cloud, previous point-based methods, such as PointNet++, extract visual features through the hierarchical aggregation of local features. However, such methods have several critical limitations: 1) They require considerable sampling and grouping operations, which leads to low inference speed. 2) Despite redundancy among adjacent points, they treat all points alike with an equal amount of computation. 3) They aggregate local features together through downsampling, which causes information loss and hurts perception capability. To overcome these challenges, we propose a novel, simple, and elegant deep learning model called YOGO (You Only Group Once). YOGO divides a point-cloud into a small number of parts and extracts a high-dimensional token to represent points within each sub-region. Next, we use self-attention to capture token-to-token relations, and project the token features back to the point features. We formulate such a series of operations as a relation inference module (RIM). Compared with previous methods, YOGO is very efficient because it only needs to sample and group a point-cloud once. Instead of operating on points, YOGO operates on a small number of tokens, each of which summarizes the point features in a sub-region. This allows us to avoid redundant computation and thus boosts efficiency. Moreover, YOGO preserves pointwise features by projecting token features to point features although the RIM computes on tokens. This avoids information loss and enhances point-wise perception capability. We conduct thorough experiments to demonstrate that YOGO achieves at least 3.0x speedup over point-based baselines while delivering competitive classification and segmentation performance on a classification dataset and a segmentation dataset based on 3D Wharehouse, and S3DIS datasets. The code is available at https://github.com/chenfengxu714/YOGO.git. Chenfeng Xu, Bohan Zhai, Bichen Wu, Peter Vajda, Kurt Keutzer, Masayoshi Tomizuka |
IROS | 8 |
| 2021 | Diverse Critical Interaction Generation for Planning and Planner EvaluationabstractGenerating diverse and comprehensive interacting agents to evaluate the decision-making modules is essential for the safe and robust planning of autonomous vehicles (AV). Due to efficiency and safety concerns, most researchers choose to train interactive adversary (competitive or weakly competitive) agents in simulators and generate test cases to interact with evaluated AVs. However, most existing methods fail to provide both natural and critical interaction behaviors in various traffic scenarios. To tackle this problem, we propose a styled generative model RouteGAN that generates diverse interactions by controlling the vehicles separately with desired styles. By altering its style coefficients, the model can generate trajectories with different safety levels serve as an online planner. Experiments show that our model can generate diverse interactions in various scenarios. We evaluate different planners with our model by testing their collision rate in interaction with RouteGAN planners of multiple critical levels. Zhao-Heng Yin, Lingfeng Sun, Masayoshi Tomizuka |
IROS | 4 |
| 2021 | Automatic Construction of Lane-level HD Maps for Urban ScenesabstractHigh definition (HD) maps have demonstrated their essential roles in enabling full autonomy, especially in complex urban scenarios. As a crucial layer of the HD map, lane-level maps are particularly useful: they contain geometrical and topological information for both lanes and intersections. However, large scale construction of HD maps is limited by tedious human labeling and high maintenance costs, especially for urban scenarios with complicated road structures and irregular markings. This paper proposes an approach based on semantic-particle filter to tackle the automatic lane-level mapping problem in urban scenes. The map skeleton is firstly structured as a directed cyclic graph from online mapping database OpenStreetMap. Our proposed method then performs semantic segmentation on 2D front-view images from ego vehicles and explores the lane semantics on a birds-eye-view domain with true topographical projection. Exploiting OpenStreetMap, we further infer lane topology and reference trajectory at intersections with the aforementioned lane semantics. The proposed algorithm has been tested in densely urbanized areas, and the results demonstrate accurate and robust reconstruction of the lane-level HD map. Yiyang Zhou, Yuichi Takeda, Masayoshi Tomizuka |
IROS | 3 |
| 2021 | Exploring Social Posterior Collapse in Variational Autoencoder for Interaction ModelingabstractMulti-agent behavior modeling and trajectory forecasting are crucial for the safe navigation of autonomous agents in interactive scenarios. Variational Autoencoder (VAE) has been widely applied in multi-agent interaction modeling to generate diverse behavior and learn a low-dimensional representation for interacting systems. However, existing literature did not formally discuss if a VAE-based model can properly encode interaction into its latent space. In this work, we argue that one of the typical formulations of VAEs in multi-agent modeling suffers from an issue we refer to as social posterior collapse, i.e., the model is prone to ignoring historical social context when predicting the future trajectory of an agent. It could cause significant prediction errors and poor generalization performance. We analyze the reason behind this under-explored phenomenon and propose several measures to tackle it. Afterward, we implement the proposed framework and experiment on real-world datasets for multi-agent trajectory prediction. In particular, we propose a novel sparse graph attention message-passing (sparse-GAMP) layer, which helps us detect social posterior collapse in our experiments. In the experiments, we verify that social posterior collapse indeed occurs. Also, the proposed measures are effective in alleviating the issue. As a result, the model attains better generalization performance when historical social context is informative for prediction. Chen Tang 0001, Masayoshi Tomizuka |
NeurIPS | 3 |
| 2021 | Neural-Network-Based Iterative Learning Control for Multiple TasksabstractIterative learning control (ILC) can synthesize the feedforward control signal for the trajectory tracking control of a repetitive task, even when the system has strong nonlinear dynamics. This makes ILC be one of the most popular methods for trajectory tracking control. Restriction on a repetitive task, however, limits its application to multiple trajectories. This article proposes a neural-network-based ILC (NN-ILC) to deal with nonrepetitive tasks very effectively. A position-based ILC is designed to compensate the tracking error, based on which the multiple outputs of the ILC (ILC outputs) for multiple tasks are expressed as a function of the reference position, velocity, and acceleration. The proposed NN-ILC divides the ILC outputs of multiple tasks into two parts: the linear and nonlinear portions. The first part is expressed by a linear function, which is the linear portion of the function of the ILC outputs. The second part is expressed by a nonlinear function, which is estimated by complementary neural networks including a general neural network and a switching neural network. Finally, the two parts are combined and the ILC outputs of multiple tasks are expressed as a neural-network-based function. Two advantages of the proposed NN-ILC are emphasized. First, the ILC outputs of multiple tasks are compressed into a function by the proposed method, and thus, the memories can be saved. Second, in terms of generalizability, the neural-network-based function of the ILC outputs can easily predict position compensation for multiple tasks without extra iterative learning processes. Experimental results on a robot arm show that the proposed NN-ILC method can easily realize the ILC of multiple tasks. It can save memory comparing with the method of storing the data of multiple tasks and can predict the ILC output of any task, which can accelerate the iterative learning process. Dailin Zhang, Masayoshi Tomizuka |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | SqueezeSegV3: Spatially-Adaptive Convolution for Efficient Point-Cloud Segmentation
Chenfeng Xu, Bichen Wu, Peter Vajda, Kurt Keutzer, Masayoshi Tomizuka |
ECCV (28) | 7 |
| 2020 | Analyzing the Suitability of Cost Functions for Explaining and Imitating Human Driving Behavior based on Inverse Reinforcement LearningabstractAutonomous vehicles are sharing the road with human drivers. In order to facilitate interactive driving and cooperative behavior in dense traffic, a thorough understanding and representation of other traffic participants' behavior are necessary. Cost functions (or reward functions) have been widely used to describe the behavior of human drivers since they can not only explicitly incorporate the rationality of human drivers and the theory of mind (TOM), but also share similarity with the motion planning problem of autonomous vehicles. Hence, more human-like driving behavior and comprehensible trajectories can be generated to enable safer interaction and cooperation. However, the selection of cost functions in different driving scenarios is not trivial, and there is no systematic summary and analysis for cost function selection and learning from a variety of driving scenarios. In this work, we aim to investigate to what extent cost functions are suitable for explaining and imitating human driving behavior. Further, we focus on how cost functions differ from each other in different driving scenarios. Towards this goal, we first comprehensively review existing cost function structures in literature. Based on that, we point out required conditions for demonstrations to be suitable for inverse reinforcement learning (IRL). Finally, we use IRL to explore suitable features and learn cost function weights from human driven trajectories in three different scenarios. Maximilian Naumann, Masayoshi Tomizuka |
ICRA | 4 |
| 2020 | Precise 3D Calibration of Wafer Handling Robot by Visual Detection and Tracking of Elliptic-shape WafersabstractThis work provides a framework for the 3D calibration of wafers and a wafer handling robot by monocular vision. The proposed method precisely reconstructs the 3D poses of wafers from a set of images captured by the camera mounted on the robot. In addition, it calibrates the robot kinematics simultaneously. A robust ellipse detection and tracking algorithm based on the edge arcs is developed to recognize wafers among images. Then a joint optimization is constructed from a multi-object pose graph to solve the 3D poses of wafers and other calibration parameters of the robot-camera system. The proposed tracking method is able to associate multiple incomplete elliptic segments using a Gaussian Mixture Model-based registration algorithm. The algorithm is point-based where no feature descriptor is required. The proposed 3D pose optimization incorporates shape constraints, and is more accurate than the point-wise reconstruction produced by classic bundle adjustment methods. Masayoshi Tomizuka |
ICRA | 2 |
| 2020 | UrbanLoco: A Full Sensor Suite Dataset for Mapping and Localization in Urban ScenesabstractMapping and localization is a critical module of autonomous driving, and significant achievements have been reached in this field. Beyond Global Navigation Satellite System (GNSS), research in point cloud registration, visual feature matching, and inertia navigation has greatly enhanced the accuracy and robustness of mapping and localization in different scenarios. However, highly urbanized scenes are still challenging: LIDAR- and camera-based methods perform poorly with numerous dynamic objects; the GNSS-based solutions experience signal loss and multi-path problems; the inertia measurement units (IMU) suffer from drifting. Unfortunately, current public datasets either do not adequately address this urban challenge or do not provide enough sensor information related to map-ping and localization. Here we present UrbanLoco: a mapping/localization dataset collected in highly-urbanized environments with a full sensor-suite. The dataset includes 13 trajectories collected in San Francisco and Hong Kong, covering a total length of over 40 kilometers. Our dataset includes a wide variety of urban terrains: urban canyons, bridges, tunnels, sharp turns, etc. More importantly, our dataset includes information from LIDAR, cameras, IMU, and GNSS receivers. Now the dataset is publicly available through the link in the footnote1. Weisong Wen, Yiyang Zhou, Guohao Zhang, Saman Fahandezh-Saadi, Xiwei Bai, Masayoshi Tomizuka, Li-Ta Hsu |
ICRA | 7 |
| 2020 | End-to-end Autonomous Driving Perception with Sequential Latent Representation LearningabstractCurrent autonomous driving systems are composed of a perception system and a decision system. Both of them are divided into multiple subsystems built up with lots of human heuristics. An end-to-end approach might clean up the system and avoid huge efforts of human engineering, as well as obtain better performance with increasing data and computation resources. Compared to the decision system, the perception system is more suitable to be designed in an end-to-end framework, since it does not require online driving exploration. In this paper, we propose a novel end-to-end approach for autonomous driving perception. A latent space is introduced to capture all relevant features useful for perception, which is learned through sequential latent representation learning. The learned end-to-end perception model is able to solve the detection, tracking, localization and mapping problems altogether with only minimum human engineering efforts and without storing any maps online. The proposed method is evaluated in a realistic urban driving simulator, with both camera image and lidar point cloud as sensor inputs. The codes and videos of this work are available at our github repo†and project website‡. Jianyu Chen 0002, Masayoshi Tomizuka |
IROS | 3 |
| 2020 | Learning-Based Controller Optimization for Repetitive Robotic TasksabstractDynamic control for robotic automation tasks is traditionally designed and optimized with a model-based approach, and the performance relies heavily upon accurate system modeling. However, modeling the true dynamics of increasingly complex robotic systems is an extremely challenging task and it often renders the automation system to operate in a non-optimal condition. Notably, many industrial robotic applications involve repetitive motions and constantly generate a large amount of motion data under the non-optimal condition. These motion data contain rich information, and therefore an intelligent automation system should be able to learn from these non-optimal motion data to drive the system to operate optimally in a data-driven manner. In this paper, we propose a learning-based controller optimization algorithm for repetitive robotic tasks. To achieve this, a multi-objective cost function is designed to take into consideration both the trajectory tracking accuracy and smoothness, and then a data-driven approach is developed to estimate the gradient and Hessian based on the motion data for optimization without relying on the dynamic model. Experiments based on a magnetically-levitated nanopositioning system are conducted to demonstrate the effectiveness and practical appeals of the proposed algorithm in repetitive robotic automation tasks. Xiaocong Li, Haiyue Zhu, Jun Ma 0008, Tat Joo Teo, Chek Sing Teo, Masayoshi Tomizuka, Tong Heng Lee |
IROS | 6 |
| 2020 | A Game-Theoretic Strategy-Aware Interaction Algorithm with Validation on Real Traffic DataabstractInteractive decision-making and motion planning are important to safety-critical autonomous agents, particularly when they interact with humans. Many different interaction strategies can be exploited by humans. For instance, they might ignore the autonomous agents, or might behave as selfish optimizers by treating the autonomous agents as opponents, or might assume themselves as leaders and the autonomous agents as followers who should take responsive actions. Different interaction strategies can lead to quite different closed-loop dynamics, and misalignment between the human's policy and the autonomous agent's belief over the policy will severely impact both safety and efficiency. Moreover, a human's interaction policy can change as interaction goes on. Hence, autonomous agents need to be aware of such uncertainties on the human policy, and integrate such information into their decision-making and motion planning algorithms. In this paper, we propose a policy-aware interaction strategy based on game theory. The goal is to allow autonomous agents to estimate humans' interactive policies and respond consequently. We validate the proposed algorithm with a roundabout scenario with real traffic data. The results show that the proposed algorithm can yield trajectories that are more similar to the ground truth than those with fixed policies. Also, we estimate how humans adjust their interaction strategies statistically based on the proposed algorithm. Mu Cai, Masayoshi Tomizuka |
IROS | 4 |
| 2020 | Expressing Diverse Human Driving Behavior with Probabilistic Rewards and Online InferenceabstractIn human-robot interaction (HRI) systems, such as autonomous vehicles, understanding and representing human behavior are important. Human behavior is naturally rich and diverse. Cost/reward learning, as an efficient way to learn and represent human behavior, has been successfully applied in many domains. Most of traditional inverse reinforcement learning (IRL) algorithms, however, cannot adequately capture the diversity of human behavior since they assume that all behavior in a given dataset is generated by a single cost function. In this paper, we propose a probabilistic IRL framework that directly learns a distribution of cost functions in continuous domain. Evaluations on both synthetic data and real human driving data are conducted. Both the quantitative and subjective results show that our proposed framework can better express diverse human driving behaviors, as well as extracting different driving styles that match what human participants interpret in our user study. Zheng Wu 0002, Hengbo Ma, Masayoshi Tomizuka |
IROS | 4 |
| 2020 | Inferring Spatial Uncertainty in Object DetectionabstractThe availability of real-world datasets is the prerequisite for developing object detection methods for autonomous driving. While ambiguity exists in object labels due to error-prone annotation process or sensor observation noises, current object detection datasets only provide deterministic annotations without considering their uncertainty. This precludes an in-depth evaluation among different object detection methods, especially for those that explicitly model predictive probability. In this work, we propose a generative model to estimate bounding box label uncertainties from LiDAR point clouds, and define a new representation of the probabilistic bounding box through spatial distribution. Comprehensive experiments show that the proposed model represents uncertainties commonly seen in driving scenarios. Based on the spatial distribution, we further propose an extension of IoU, called the Jaccard IoU (JIoU), as a new evaluation metric that incorporates label uncertainty. Experiments on the KITTI and the Waymo Open Datasets show that JIoU is superior to IoU when evaluating probabilistic object detectors. Di Feng, Yiyang Zhou, Lars Rosenbaum, Fabian Timm, Klaus Dietmayer, Masayoshi Tomizuka |
IROS | 7 |
| 2020 | Application Specific System Identification for Model-Based Control in Self-Driving CarsabstractLinear Parameter Varying (LPV) models can be used to describe the vehicular lateral dynamic behavior of self-driving cars. They are particularly suitable for model-based control schemes such as model predictive control (MPC) applied to real-time trajectory tracking control, since they provide a proper trade-off between accuracy in different scenarios and reduced computation cost compared to nonlinear models. The MPC control schemes use the model for a long prediction horizon of the states, therefore prediction errors for a long time horizon should be minimized in order to increase the accuracy of the tracking. For this task, this work presents a system identification procedure for the lateral dynamics of a vehicle that combines a LPV model with a learning algorithm that has been successfully applied to other dynamic systems in the past. Simulation results show the benefits of the identified model in comparison to other well-known vehicular lateral dynamic models. Julian M. Salt Ducaju, Chen Tang 0001, Masayoshi Tomizuka, Ching-Yao Chan |
IV | 3 |
| 2020 | epBRM: Improving a Quality of 3D Object Detection using End Point Box Regression ModuleabstractWe present an endpoint box regression module(epBRM), which is designed for predicting precise 3D bounding boxes using raw LiDAR 3D point clouds. The proposed epBRM is built with sequence of small networks and is computationally lightweight. Our approach can improve a 3D object detection performance by predicting more precise 3D bounding box coordinates. The proposed approach requires 40 minutes of training to improve the detection performance. Moreover, epBRM imposes less than 12ms to network inference time for up-to 20 objects. The proposed approach utilizes a spatial transformation mechanism to simplify the box regression task. Adopting spatial transformation mechanism into epBRM makes it possible to improve the quality of detection with a small sized network. We conduct in-depth analysis of the effect of various spatial transformation mechanisms applied on raw LiDAR 3D point clouds. We also evaluate the proposed epBRM by applying it to several state-of-the-art 3D object detection systems. We evaluate our approach on KITTI dataset[1], a standard 3D object detection benchmark for autonomous vehicles. The proposed epBRM enhances the overlaps between ground truth bounding boxes and detected bounding boxes, and improves 3D object detection. Our proposed method evaluated in KITTI test server outperforms current state-of-the-art approaches. Kiwoo Shin, Masayoshi Tomizuka |
IV | 2 |
| 2020 | EvolveGraph: Multi-Agent Trajectory Prediction with Dynamic Relational ReasoningabstractMulti-agent interacting systems are prevalent in the world, from purely physical systems to complicated social dynamic systems. In many applications, effective understanding of the situation and accurate trajectory prediction of interactive agents play a significant role in downstream tasks, such as decision making and planning. In this paper, we propose a generic trajectory forecasting framework (named EvolveGraph) with explicit relational structure recognition and prediction via latent interaction graphs among multiple heterogeneous, interactive agents. Considering the uncertainty of future behaviors, the model is designed to provide multi-modal prediction hypotheses. Since the underlying interactions may evolve even with abrupt changes, and different modalities of evolution may lead to different outcomes, we address the necessity of dynamic relational reasoning and adaptively evolving the interaction graphs. We also introduce a double-stage training pipeline which not only improves training efficiency and accelerates convergence, but also enhances model performance. The proposed framework is evaluated on both synthetic physics simulations and multiple real-world benchmark datasets in various areas. The experimental results illustrate that our approach achieves state-of-the-art performance in terms of prediction accuracy. Jiachen Li 0001, Masayoshi Tomizuka, Chiho Choi |
NeurIPS | 3 |
| 2020 | Generic Tracking and Probabilistic Prediction Framework and Its Application in Autonomous DrivingabstractAccurately tracking and predicting behaviors of surrounding objects are key prerequisites for intelligent systems such as autonomous vehicles to achieve safe and high-quality decision making and motion planning. However, there still remain challenges for multi-target tracking due to object number fluctuation and occlusion. To overcome these challenges, we propose a constrained mixture sequential Monte Carlo (CMSMC) method in which a mixture representation is incorporated in the estimated posterior distribution to maintain multi-modality. Multiple targets can be tracked simultaneously within a unified framework without explicit data association between observations and tracking targets. The framework can incorporate an arbitrary prediction model as the implicit proposal distribution of the CMSMC method. An example in this paper is a learning-based model for hierarchical time-series prediction, which consists of a behavior recognition module and a state evolution module. Both modules in the proposed model are generic and flexible so as to be applied to a class of time-series prediction problems where behaviors can be separated into different levels. Finally, the proposed framework is applied to a numerical case study as well as a task of on-road vehicle tracking, behavior recognition, and prediction in highway scenarios. Instead of only focusing on forecasting trajectory of a single entity, we jointly predict continuous motions for interactive entities simultaneously. The proposed approaches are evaluated from multiple aspects, which demonstrate great potential for intelligent vehicular systems and traffic surveillance systems. Jiachen Li 0001, Yeping Hu, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Disturbance-Observer-Based Tracking Controller for Neural Network Driving Policy TransferabstractThe neural network policies are widely explored in the autonomous driving field, thanks to their capability of handling complicated driving tasks. However, the practical deployment of such policies is slowed down due to their lack of robustness against modeling gap and external disturbances. In our prior work, we proposed a planner-controller architecture and applied a disturbance-observer-based (DOB) robust tracking controller to reject the disturbances and achieved zero-shot policy transfer. In this paper, we present our latest progress on improving the policy transfer performance under this framework. Concretely, we applied adaptive DOB, so as to more accurately model the inverse system dynamics and increase the cut-off frequency of the Q-filter in the DOB. A closed-loop reference path smoothing algorithm is introduced to alleviate the step disturbance input imposed by the reference trajectory re-planning. On the neural network control policy side, we applied the parallel attribute networks, a hierarchical modular policy network to dynamically handle various driving tasks. We have carried out various simulations and experiments to validate the capability of our proposed method to achieve sim-to-sim and sim-to-real policy transfer. The proposed method achieves the most outstanding performance among a series of baseline control schemes. Chen Tang 0001, Masayoshi Tomizuka |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | A Learning Framework for High Precision Industrial AssemblyabstractAutomatic assembly has broad applications in industries. Traditional assembly tasks utilize predefined trajectories or tuned force control parameters, which make the automatic assembly time-consuming, difficult to generalize, and not robust to uncertainties. In this paper, we propose a learning framework for high precision industrial assembly. The framework combines both the supervised learning and the reinforcement learning. The supervised learning utilizes trajectory optimization to provide the initial guidance to the policy, while the reinforcement learning utilizes actor-critic algorithm to establish the evaluation system even the supervisor is not accurate. The proposed learning framework is more efficient compared with the reinforcement learning and achieves better stability performance than the supervised learning. The effectiveness of the method is verified by both the simulation and experiment. Experimental videos are available at [1]. Yongxiang Fan, Jieliang Luo, Masayoshi Tomizuka |
ICRA | 3 |
| 2019 | Interaction-aware Multi-agent Tracking and Probabilistic Behavior Prediction via Adversarial LearningabstractIn order to enable high-quality decision making and motion planning of intelligent systems such as robotics and autonomous vehicles, accurate probabilistic predictions for surrounding interactive objects is a crucial prerequisite. Although many research studies have been devoted to making predictions on a single entity, it remains an open challenge to forecast future behaviors for multiple interactive agents simultaneously. In this work, we take advantage of the Generative Adversarial Network (GAN) due to its capability of distribution learning and propose a generic multi-agent probabilistic prediction and tracking framework which takes the interactions among multiple entities into account, in which all the entities are treated as a whole. However, since GAN is very hard to train, we make an empirical research and present the relationship between training performance and hyperparameter values with a numerical case study. The results imply that the proposed model can capture both the mean, variance and multi-modalities of the groundtruth distribution. Moreover, we apply the proposed approach to a real-world task of vehicle behavior prediction to demonstrate its effectiveness and accuracy. The results illustrate that the proposed model trained by adversarial learning can achieve a better prediction performance than other state-of-the-art models trained by traditional supervised learning which maximizes the data likelihood. The well-trained model can also be utilized as an implicit proposal distribution for particle filtered based Bayesian state estimation. Jiachen Li 0001, Hengbo Ma, Masayoshi Tomizuka |
ICRA | 3 |
| 2019 | Adaptive Probabilistic Vehicle Trajectory Prediction Through Physically Feasible Bayesian Recurrent Neural NetworkabstractProbabilistic vehicle trajectory prediction is essential for robust safety of autonomous driving. Current methods for long-term trajectory prediction cannot guarantee the physical feasibility of predicted distribution. Moreover, their models cannot adapt to the driving policy of the predicted target human driver. In this work, we propose to overcome these two shortcomings by a Bayesian recurrent neural network model consisting of Bayesian-neural-network-based policy model and known physical model of the scenario. Bayesian neural network can ensemble complicated output distribution, enabling rich family of trajectory distribution. The embedded physical model ensures feasibility of the distribution. Moreover, the adopted gradient-based training method allows direct optimization for better performance in long prediction horizon. Furthermore, a particle-filter-based parameter adaptation algorithm is designed to adapt the policy Bayesian neural network to the predicted target online. Effectiveness of the proposed methods is verified with a toy example with multi-modal stochastic feedback gain and naturalistic car following data. Chen Tang 0001, Jianyu Chen 0002, Masayoshi Tomizuka |
ICRA | 3 |
| 2019 | Deep Imitation Learning for Autonomous Driving in Generic Urban Scenarios with Enhanced SafetyabstractThe decision and planning system for autonomous driving in urban environments is hard to design. Most current methods manually design the driving policy, which can be expensive to develop and maintain at scale. Instead, with imitation learning we only need to collect data and the computer will learn and improve the driving policy automatically. However, existing imitation learning methods for autonomous driving are hardly performing well for complex urban scenarios. Moreover, the safety is not guaranteed when we use a deep neural network policy. In this paper, we proposed a framework to learn the driving policy in urban scenarios efficiently given offline connected driving data, with a safety controller incorporated to guarantee safety at test time. The experiments show that our method can achieve high performance in realistic simulations of urban driving scenarios. Jianyu Chen 0002, Bodi Yuan, Masayoshi Tomizuka |
IROS | 3 |
| 2019 | optimization Model for Planning Precision Grasps with Multi-Fingered HandsabstractPrecision grasps with multi-fingered hands are important for precise placement and in-hand manipulation tasks. Searching precision grasps on the object represented by point cloud, is challenging due to the complex object shape, high-dimensionality, collision and undesired properties of the sensing and positioning. This paper proposes an optimization model to search for precision grasps with multi-fingered hands. The model takes noisy point cloud of the object as input and optimizes the grasp quality by iteratively searching for the palm pose and finger joints positions. The collision between the hand and the object is approximated and penalized by a series of least-squares. The collision approximation is able to handle the point cloud representation of the objects with complex shapes. The proposed optimization model is able to locate collision-free optimal precision grasps efficiently. The average computation time is 0.50 sec/grasp. The searching is robust to the incompleteness and noise of the point cloud. The effectiveness of the algorithm is demonstrated by experiments. Yongxiang Fan, Xinghao Zhu, Masayoshi Tomizuka |
IROS | 3 |
| 2019 | Interaction-aware Decision Making with Adaptive Strategies under Merging ScenariosabstractIn order to drive safely and efficiently under merging scenarios, autonomous vehicles should be aware of their surroundings and make decisions by interacting with other road participants. Moreover, different strategies should be made when the autonomous vehicle is interacting with drivers having different level of cooperativeness. Whether the vehicle is on the merge-lane or main-lane will also influence the driving maneuvers since drivers will behave differently when they have the right-of-way than otherwise. Many traditional methods have been proposed to solve decision making problems under merging scenarios. However, these works either are incapable of modeling complicated interactions or require implementing hand-designed rules which cannot properly handle the uncertainties in real-world scenarios. In this paper, we proposed an interaction-aware decision making with adaptive strategies (IDAS) approach that can let the autonomous vehicle negotiate the road with other drivers by leveraging their cooperativeness under merging scenarios. A single policy is learned under the multi-agent reinforcement learning (MARL) setting via the curriculum learning strategy, which enables the agent to automatically infer other drivers' various behaviors and make decisions strategically. A masking mechanism is also proposed to prevent the agent from exploring states that violate common sense of human judgment and increase the learning efficiency. An exemplar merging scenario was used to implement and examine the proposed method. Yeping Hu, Alireza Nakhaei, Masayoshi Tomizuka, Kikuo Fujimura |
IROS | 3 |
| 2019 | Robust Deformation Model Approximation for Robotic Cable ManipulationabstractCable manipulation is a challenging task for robots. The major challenge is that cables have high degrees of freedom and are easy to deform during manipulation. In this paper, we propose a novel framework SPR-RWLS to manipulate cables, which includes real-time cable tracking and robust local deformation model approximation. For cable tracking, structure preserved registration (SPR) is utilized to robustly estimate the movement of selected points on a cable even in the presence of sensor noise, outliers, and occlusions. Robust weighted least squares (RWLS) is then applied to calculate the local deformation model of the cable under uncertainties. We show that SPR-RWLS enables the dual-arm robots to manipulate cables with different thicknesses and lengths to different desired curvatures in multiple scenarios. We also show that real-time implementation of the proposed method can be simplified by parallel computation. Shiyu Jin, Masayoshi Tomizuka |
IROS | 3 |
| 2019 | Prediction of Human Arm Target for Robot Reaching MovementsabstractThe raise of collaborative robotics has allowed to create new spaces where robots and humans work in proximity. Consequently, to predict human movements and his/her final intention becomes crucial to anticipate robot next move, preserving safety and increasing efficiency. In this paper we propose a human-arm prediction algorithm that allows to infer if the human operator is moving towards the robot to intentionally interact with it. The human hand position is tracked by an RGB-D camera online. By combining the Minimum Jerk model with Semi-Adaptable Neural Networks we obtain a reliable prediction of the human hand trajectory and final target in a short amount of time. The proposed algorithm was tested in a multi-movements scenario with FANUC LR Mate 200iD/7L industrial robot. Chiara Talignani Landi, Yujiao Cheng, Federica Ferraguti, Marcello Bonfè, Cristian Secchi, Masayoshi Tomizuka |
IROS | 6 |
| 2019 | Conditional Generative Neural System for Probabilistic Trajectory PredictionabstractEffective understanding of the environment and accurate trajectory prediction of surrounding dynamic obstacles are critical for intelligent systems such as autonomous vehicles and wheeled mobile robotics navigating in complex scenarios to achieve safe and high-quality decision making, motion planning and control. Due to the uncertain nature of the future, it is desired to make inference from a probability perspective instead of deterministic prediction. In this paper, we propose a conditional generative neural system (CGNS) for probabilistic trajectory prediction to approximate the data distribution, with which realistic, feasible and diverse future trajectory hypotheses can be sampled. The system combines the strengths of conditional latent space learning and variational divergence minimization, and leverages both static context and interaction information with soft attention mechanisms. We also propose a regularization method for incorporating soft constraints into deep neural networks with differentiable barrier functions, which can regulate and push the generated samples into the feasible regions. The proposed system is evaluated on several public benchmark datasets for pedestrian trajectory prediction and a roundabout naturalistic driving dataset collected by ourselves. The experimental results demonstrate that our model achieves better performance than various baseline approaches in terms of prediction accuracy. Jiachen Li 0001, Hengbo Ma, Masayoshi Tomizuka |
IROS | 3 |
| 2019 | Precise Correntropy-based 3D Object Modelling With Geometrical Traffic PriorabstractRobust 3D perception using LiDAR is of prime importance for robotics, and its fundamental core lies in precise object modelling resisting to noise and outliers. In this paper, a precise 3D object modelling algorithm is designed especially for the intelligent vehicles. The proposed algorithm is advantageous by leveraging the crucial traffic geometrical prior of road surface profile, and both the noise and outliers are elegantly handled by robust correntropy-based metric. More specifically, the road surface correction (RSC) method transforms each individual LiDAR measurement from its locally planar road surface to a globally ideal plane. This procedure essentially guarantees the reduction of vehicle's motion from arbitrary 3D motion to physically feasible 2D motion. To deal with the noise and outliers, a correntropy-based multi-frame matching (CorrMM) algorithm is proposed which has a robust objective function with respect to point-to-plane residual error. An efficient solver inspired by M-estimator and retraction technique on Lie group is developed, which elegantly converts the optimization of highly non-linear objective function into a simple quadratic programming (QP) problem. Extensive experimental results validate that the proposed algorithm attains more crisper 3D object models than several state-of-the-art algorithms on a challenging real traffic dataset. Di Wang 0028, Jianru Xue, Yinghan Jin, Nanning Zheng 0001, Masayoshi Tomizuka |
IROS | 6 |
| 2019 | Constructing a Highly Interactive Vehicle Motion DatasetabstractResearch in the areas related to driving behavior, e.g., behavior modeling and prediction, requires datasets with highly interactive vehicle motions. Existing public vehicle motion datasets emphasize increasing the number of vehicles and time duration, but behavior-related researchers are suffering from two factors. First, strong interactions among vehicles are not well addressed and datasets are of relatively low-density to observe meaningful interactions. Second, most of the existing datasets are missing the map information with reference paths which is essential for driving-behavior-related research. To address this issue, a dataset with highly interactive vehicle motions is constructed in this paper. A variety of challenging driving scenarios such as unsignalized intersections and roundabouts are included. Reference paths are also constructed from motion data along with high-definition maps so that key features can be generated for both prediction and planning algorithms. Moreover, we propose a set of metrics to extract the interactive motions in different maps, including the minimum difference of time to collision point (MDTTC) and duration of waiting period. Such metrics are used to quantity the interaction density of the dataset. We also give several representative results on prediction and motion generation utilizing the constructed dataset to demonstrate how the dataset can facilitate research in the area of driving behavior. Di Wang 0028, Yinghan Jin, Masayoshi Tomizuka |
IROS | 5 |
| 2019 | Multi-modal Probabilistic Prediction of Interactive Behavior via an Interpretable ModelabstractFor autonomous agents to successfully operate in real world, the ability to anticipate future motions of surrounding entities in the scene can greatly enhance their safety levels since potentially dangerous situations could be avoided in advance. While impressive results have been shown on predicting each agent's behavior independently, we argue that it is not valid to consider road entities individually since transitions of vehicle states are highly coupled. Moreover, as the predicted horizon becomes longer, modeling prediction uncertainties and multi-modal distributions over future sequences will turn into a more challenging task. In this paper, we address this challenge by presenting a multi-modal probabilistic prediction approach. The proposed method is based on a generative model and is capable of jointly predicting sequential motions of each pair of interacting agents. Most importantly, our model is interpretable, which can explain the underneath logic as well as obtain more reliability to use in real applications. A complicate real-world roundabout scenario is utilized to implement and examine the proposed method. Yeping Hu, Masayoshi Tomizuka |
IV | 4 |
| 2019 | Coordination and Trajectory Prediction for Vehicle Interactions via Bayesian Generative ModelingabstractCoordination recognition and subtle pattern prediction of future trajectories play a significant role when modeling interactive behaviors of multiple agents. Due to the essential property of uncertainty in the future evolution, deterministic predictors are not sufficiently safe and robust. In order to tackle the task of probabilistic prediction for multiple, interactive entities, we propose a coordination and trajectory prediction system (CTPS), which has a hierarchical structure including a macro-level coordination recognition module and a micro-level subtle pattern prediction module which solves a probabilistic generation task. We illustrate two types of representation of the coordination variable: categorized and real-valued, and compare their effects and advantages based on empirical studies. We also bring the ideas of Bayesian deep learning into deep generative models to generate diversified prediction hypotheses. The proposed system is tested on multi-ple driving datasets in various traffic scenarios, which achieves better performance than baseline approaches in terms of a set of evaluation metrics. The results also show that using categorized coordination can better capture multi-modality and generate more diversified samples than the real-valued coordination, while the latter can generate prediction hypotheses with smaller errors with a sacrifice of sample diversity. Moreover, employing neural networks with weight uncertainty is able to generate samples with larger variance and diversity. Jiachen Li 0001, Hengbo Ma, Masayoshi Tomizuka |
IV | 4 |
| 2019 | Wasserstein Generative Learning with Kinematic Constraints for Probabilistic Interactive Driving Behavior PredictionabstractSince prediction plays a significant role in enhancing the performance of decision making and planning procedures, the requirement of advanced methods of prediction becomes urgent. Although many literatures propose methods to make prediction on a single agent, there is still a challenging and open problem on how to make prediction for multi-agent systems. In this work, by leveraging the power of statistics and information theory, we propose a novel deep latent variable model based on Wasserstein auto-encoder, which is able to learn a complex probabilistic distribution. Models such as neural networks cannot guarantee the satisfaction of dynamic system constraints directly. Therefore, we also propose a novel generative model structure to enable our approach to satisfy the kinematic constraints automatically. We test our model on both numerical examples and a real-world application to demonstrate its accuracy and efficiency. The results show that the proposed model achieves a better prediction accuracy than the other state-of-the-art methods under common evaluation metrics. Moreover, we introduce statistics to evaluate if the generative model literally learns the interaction patterns between different agents in the environments. Hengbo Ma, Jiachen Li 0001, Masayoshi Tomizuka |
IV | 4 |
| 2019 | RoarNet: A Robust 3D Object Detection based on RegiOn Approximation RefinementabstractWe present RoarNet, a new approach for 3D object detection from 2D image and 3D Lidar point clouds. Based on two stage object detection framework ([1], [2]) with PointNet [3] as our backbone network, we suggest several novel ideas to improve 3D object detection performance. The first part of our method, RoarNet_2D, estimates the 3D poses of objects from a monocular image, which approximates where to examine further, and derives multiple candidates that are geometrically feasible. This step significantly narrows down feasible 3D regions, which otherwise requires demanding processing of 3D point clouds in a huge search space. Then the second part, RoarNet_3D, takes the candidate regions and conducts in-depth inferences to conclude final poses in a recursive manner. Inspired by PointNet, RoarNet_3D processes 3D point clouds directly without any loss of data, leading to precise detection. We evaluate our method in KITTI, a 3D object detection benchmark. Our result shows that RoarNet has superior performance to state-of-the-art methods that are publicly available. Remarkably, RoarNet also outperforms state-of-the-art methods even in settings where Lidar and camera are not time synchronized, which is practically important for actual driving environment. Kiwoo Shin, Youngwook Paul Kwon, Masayoshi Tomizuka |
IV | 3 |
| 2019 | Behavior Planning of Autonomous Cars with Social PerceptionabstractAutonomous cars have to navigate in dynamic environment which can be full of uncertainties. The uncertainties can come either from sensor limitations such as occlusions and limited sensor range, or from probabilistic prediction of other road participants, or from unknown social behavior in a new area. To safely and efficiently drive in the presence of these uncertainties, the decision-making and planning modules of autonomous cars should intelligently utilize all available information and appropriately tackle the uncertainties so that proper driving strategies can be generated. In this paper, we propose a social perception scheme which treats all road participants as distributed sensors in a sensor network. By observing the individual behaviors as well as the group behaviors, uncertainties of the three types can be updated uniformly in a belief space. The updated beliefs from the social perception are then explicitly incorporated into a probabilistic planning framework based on Model Predictive Control (MPC). The cost function of the MPC is learned via inverse reinforcement learning (IRL). Such an integrated probabilistic planning module with socially enhanced perception enables the autonomous vehicles to generate behaviors which are defensive but not overly conservative, and socially compatible. The effectiveness of the proposed framework is verified in simulation on an representative scenario with sensor occlusions. Ching-Yao Chan, Masayoshi Tomizuka |
IV | 4 |
| 2019 | Toward Modularization of Neural Network Autonomous Driving Policy Using Parallel Attribute NetworksabstractNeural network autonomous driving policies are widely explored. However, no matter using imitation learning or reinforcement learning, the network policies are generally hard to train, and the learned knowledge encoded in neural network policies are hard to transfer. We propose to modularize the complicated driving policies in terms of the driving attributes, and present the parallel attribute networks (PAN), which can learn to fullfill the requirements of the attributes in the driving tasks separately, and later assemble their knowledge together. Concretely, we first train a policy network that accomplish the base lane tracking attribute. The modules for the add-on attributes such as avoiding obstacles and obeying traffic rules are then trained to map the corresponding state to a satisfactory set of the vehicle action space. Finally the reference action given by the base policy is projected into the satisfactory sets so as to satisfy the requirements of all the attributes. Using the PAN, many complicated tasks that are hard to train from scratch can be easily trained; also unseen driving tasks can be solved in a zero-shot manner by assembling the pretrained attribute modules. We have validated the capability of our model on a class of autonomous driving problems with attributes of obstacle avoidance, traffic light and speed limit in simulation. Experimental results based on an obstacle avoidance task are also presented. Haonan Chang, Chen Tang 0001, Changliu Liu, Masayoshi Tomizuka |
IV | 5 |
| 2018 | Fast Robot Motion Planning with Collision Avoidance and Temporal OptimizationabstractConsidering the growing demand of real-time motion planning in robot applications, this paper proposes a fast robot motion planner (FRMP) to plan collision-free and time-optimal trajectories, which applies the convex feasible set algorithm (CFS) to solve both the trajectory planning problem and the temporal optimization problem. The performance of CFS in trajectory planning is compared to the sequential quadratic programming (SQP) in simulation, which shows a significant decrease in iteration numbers and computation time to converge a solution. The effectiveness of temporal optimization is shown on the operational time reduction in the experiment on FANUC LR Mate 200iD/7L. Hsien-Chung Lin, Changliu Liu, Masayoshi Tomizuka |
ICARCV | 3 |
| 2018 | Efficient Trajectory Optimization for Robot Motion PlanningabstractMotion planning for multi-jointed robots is challenging. Due to the inherent complexity of the problem, most existing works decompose motion planning as easier subproblems. However, because of the inconsistent performance metrics, only sub-optimal solution can be found by decomposition based approaches. This paper presents an optimal control based approach to address the path planning and trajectory planning subproblems simultaneously. Unlike similar works which either ignore robot dynamics or require long computation time, an efficient numerical method for trajectory optimization is presented in this paper for motion planning involving complicated robot dynamics. The efficiency and effectiveness of the proposed approach is shown by numerical results. Experimental results are used to show the feasibility of the presented planning algorithm. Yu Zhao 0015, Hsien-Chung Lin, Masayoshi Tomizuka |
ICARCV | 3 |
| 2018 | Real-Time Grasp Planning for Multi-Fingered Hands by Finger SplittingabstractGrasp planning for multi-fingered hands is computationally expensive due to the joint-contact coupling, surface nonlinearities and high dimensionality, thus is generally not affordable for real-time implementations. Traditional planning methods by optimization, sampling or learning work well in planning for parallel grippers but remain challenging for multi-fingered hands. This paper proposes a strategy called finger splitting, to plan precision grasps for multi-fingered hands starting from optimal parallel grasps. The finger splitting is optimized by a dual-stage iterative optimization including a contact point optimization (CPO) and a palm pose optimization (PPO), to gradually split fingers and adjust both the contact points and the palm pose. The dual-stage optimization is able to consider both the object grasp quality and hand manipulability, address the nonlinearities and coupling, and achieve efficient convergence within one second. Simulation results demonstrate the effectiveness of the proposed approach. The simulation video is available at [1]. Yongxiang Fan, Te Tang, Hsien-Chung Lin, Masayoshi Tomizuka |
IROS | 4 |
| 2018 | Characterization of Active/Passive Pneumatic Actuators for Assistive DevicesabstractAssistive devices have been developed for power augmentation and task-oriented assistance such as loaded walking. The effective joint dynamics of the user can be altered using a wearable system, providing assistance when a task is performed. The authors have investigated an Active/Passive Pneumatic Actuator (AP2A) for an assistive device, which has a simple structure and responds as a passive nonlinear spring with controllable stiffness. This paper introduces a novel controller for the AP2 A and validates the performance through experiments. The developed controller is found to stabilize at the desired stiffness response within 1 second, confirming the ability of the AP2 A to act as an adjustable passive nonlinear spring. Daisuke Kaneishi, Masayoshi Tomizuka, Robert Peter Matthew |
IROS | 2 |
| 2018 | A Framework for Robot Grasp Transferring with Non-rigid TransformationabstractGrasp planning is essential for robots to execute dexterous tasks. Solving the optimal grasps for various objects online, however, is challenging due to the heavy computation load during exhaustive sampling, and the difficulties to consider task requirements. This paper proposes a framework to combine analytic approach with learning for efficient grasp generation. The example grasps are taught by human demonstration and mapped to similar objects by a non-rigid transformation. The mapped grasps are evaluated analytically and refined by an orientation search to improve the grasp robustness and robot reachability. The proposed approach is able to plan high-quality grasps, avoid collision, satisfy task requirements, and achieve efficient online planning. The effectiveness of the proposed method is verified by a series of experiments. Hsien-Chung Lin, Te Tang, Yongxiang Fan, Masayoshi Tomizuka |
IROS | 4 |
| 2018 | Courteous Autonomous CarsabstractTypically, autonomous cars optimize for a combination of safety, efficiency, and driving quality. But as we get better at this optimization, we start seeing behavior go from too conservative to too aggressive. The car's behavior exposes the incentives we provide in its cost function. In this work, we argue for cars that are not optimizing a purely selfish cost, but also try to be courteous to other interactive drivers. We formalize courtesy as a term in the objective that measures the increase in another driver's cost induced by the autonomous car's behavior. Such a courtesy term enables the robot car to be aware of possible irrationality of the human behavior, and plan accordingly. We analyze the effect of courtesy in a variety of scenarios. We find, for example, that courteous robot cars leave more space when merging in front of a human driver. Moreover, we find that such a courtesy term can help explain real human driver behavior on the NGSIM dataset. Masayoshi Tomizuka, Anca D. Dragan |
IROS | 3 |
| 2018 | Continuous Decision Making for On-road Autonomous Driving under Uncertain and Interactive EnvironmentsabstractAlthough autonomous driving techniques have achieved great improvements, challenges still exist in decision making for variety of different scenarios under uncertain and interactive environments. A good decision maker must satisfy the following requirements: (1) Be in a generic and unified form to cover as more scenarios as possible. (2) Be able to interact properly with other moving obstacles under the uncertainty of their motions. In this paper, the continuous decision making (CDM) framework is proposed to formulate different driving scenarios in a unified way, which encodes the high level decision making information into a continuous reference trajectory that can be naturally combined with a lower level trajectory planner. Within the framework, a maximum interaction defensive policy (MIDP) is proposed, which calculates the best action to interact with stochastic moving obstacles while guaranteeing safety. The method is applied to a ramp merging scenario and the stochastic behavior models of the surrounding vehicles are learned from the NGSIM dataset. Simulations are shown to visualize and analyze the results. Jianyu Chen 0002, Chen Tang 0001, Long Xin, Shengbo Eben Li, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 5 |
| 2018 | Deep Hierarchical Reinforcement Learning for Autonomous Driving with Distinct BehaviorsabstractDeep reinforcement learning has achieved great progress recently in domains such as learning to play Atari games from raw pixel input. The model-free characteristics of reinforcement learning free us from hand-encoding complex policies. However, for real world tasks such as autonomous driving, there are some complex sequential decision making processes that contain distinct behaviors. Due to the delayed rewards and the averaged gradient, it is pretty difficult for a flat deep reinforcement learning algorithm to learn a good policy. In this paper, we design a hierarchical neural network policy and propose a hierarchical policy gradient method to train the network with the semi markov decision process (SMDP) temporal abstraction formulation. We apply this method to a traffic light passing scenario in autonomous driving, where the vehicle has two distinct behaviors (e.g., pass and stop) and its primitive actions (e.g., acceleration) should follow the corresponding behavior. We show via simulation that our method is able to select correct decision and acts appropriately when the traffic light turns yellow. On the contrary, the flat reinforcement learning algorithm is not able to achieve a good performance and exhibits a large variance. Furthermore, the trained neural network modules are reusable in the future to cover more scenarios. Jianyu Chen 0002, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 3 |
| 2018 | Probabilistic Prediction of Vehicle Semantic Intention and MotionabstractAccurately predicting the possible behaviors of traffic participants is an essential capability for future autonomous vehicles. The majority of current researches fix the number of driving intentions by considering only a specific scenario. However, distinct driving environments usually contain various possible driving maneuvers. Therefore, a intention prediction method that can adapt to different traffic scenarios is needed. To further improve the overall vehicle prediction performance, motion information is usually incorporated with classified intentions. As suggested in some literature, the methods that directly predict possible goal locations can achieve better performance for long-term motion prediction than other approaches due to their automatic incorporation of environment constraints. Moreover, by obtaining the temporal information of the predicted destinations, the optimal trajectories for predicted vehicles as well as the desirable path for ego autonomous vehicle could be easily generated. In this paper, we propose a Semantic based Intention and Motion Prediction (SIMP) method, which can be adapted to any driving scenarios by using semantic defined vehicle behaviors. It utilizes a probabilistic framework based on deep neural network to estimate the intentions, final locations, and the corresponding time information for surrounding vehicles. An exemplar real-world scenario was used to implement and examine the proposed method. Yeping Hu, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 3 |
| 2018 | Generic Vehicle Tracking Framework Capable of Handling Occlusions Based on Modified Mixture Particle FilterabstractAccurate and robust tracking of surrounding road participants plays an important role in autonomous driving. However, there is usually no prior knowledge of the number of tracking targets due to object emergence, object disappearance and false alarms. To overcome this challenge, we propose a generic vehicle tracking framework based on modified mixture particle filter, which can make the number of tracking targets adaptive to real-time observations and track all the vehicles within sensor range simultaneously in a uniform architecture without explicit data association. Each object corresponds to a mixture component whose distribution is non-parametric and approximated by particle hypotheses. Most tracking approaches employ vehicle kinematic models as the prediction model. However, it is hard for these models to make proper predictions when sensor measurements are lost or become low quality due to partial or complete occlusions. Moreover, these models are incapable of forecasting sudden maneuvers. To address these problems, we propose to incorporate learning-based behavioral models instead of pure vehicle kinematic models to realize prediction in the prior update of recursive Bayesian state estimation. Two typical driving scenarios including lane keeping and lane change are demonstrated to verify the effectiveness and accuracy of the proposed framework as well as the advantages of employing learning-based models. Jiachen Li 0001, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 3 |
| 2018 | Cooperative Driving Based on Negotiation with Persuasion and ConcessionabstractRecently, along with emergence of autonomous driving vehicles, it is predicted that in the near future, human drivers need to share road and interact with self-driving cars in all the possible traffic scenarios. In this paper, an algorithm based on negotiation with both persuasion and concession is proposed to tackle the challenging task of cooperative driving involving both human and robot drivers. The decision making process is formulated as an optimization based negotiation problem. The persuasion of autonomous vehicle is achieved by making commitment to tipping towards cooperation in the form of convex constraint. The concession is accomplished by gradually tuning weights in the objective function. We propose an approach suitable for most common driving scenarios including ramp merging, lane keeping/changing and intersection crossing. The effectiveness of the proposed algorithm is demonstrated by simulation for several different driving scenarios. Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 2 |
| 2018 | Fusing Bird's Eye View LIDAR Point Cloud and Front View Camera Image for 3D Object DetectionabstractWe propose a new method for fusing LIDAR point cloud and camera-captured images in deep convolutional neural networks (CNN). The proposed method constructs a new layer called sparse non-homogeneous pooling layer to transform features between bird's eye view and front view. The sparse point cloud is used to construct the mapping between the two views. The pooling layer allows efficient fusion of the multi-view features at any stage of the network. This is favorable for 3D object detection using camera-LIDAR fusion for autonomous driving. A corresponding one-stage detector is designed and tested on the KITTI bird's eye view object detection dataset, which produces 3D bounding boxes from the bird's eye view map. The fusion method shows significant improvement on both speed and accuracy of the pedestrian detection over other fusion-based object detection networks. Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 3 |
| 2018 | Probabilistic Prediction from Planning Perspective: Problem Formulation, Representation Simplification and Evaluation MetricabstractAccurate probabilistic prediction for intention and motion of road users is a key prerequisite to achieve safe and high-quality decision-making and motion planning for autonomous driving. Typically, the performance of probabilistic predictions was only evaluated by learning metrics for approximation to the motion distribution in the dataset. However, as a module supporting decision and planning, probabilistic prediction should also be evaluated from decision and planning perspective. Moreover, the evaluation of probabilistic prediction highly relies on the problem formulation variation and motion representation simplification, which lacks a formal foundation in a comprehensive framework. To address such concerns, we provide a systematic and unified framework for the analysis of three under-explored aspects of probabilistic prediction: problem formulation, representation simplification and evaluation metric. More importantly, we address the omitted but crucial problems in the three aspects from decision and planning perspective. In addition to a review of learning metrics, metrics to be considered from planning perspective are highlighted, such as planning consequence of inaccurate and erroneous prediction, as well as violations of predicted motions to planning constraints. We address practical formulation variations of prediction problems, such as decision-maker view and blind view for viewpoint, as well as reactive prediction for interaction, so that decision and planning can be facilitated. Arnaud de La Fortelle, Yi-Ting Chen 0001, Ching-Yao Chan, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 5 |
| 2018 | Non-uniform Multi-rate Estimator based Periodic Event-Triggered Control for resource saving
Ángel Cuenca, Minghui Zheng, Masayoshi Tomizuka, Sergio Sánchez |
Inf. Sci. | 3 |
| 2017 | Safe and feasible motion generation for autonomous driving via constrained policy netabstractPolicy networks have great potential to learn sophisticated driving policy under complicated interaction between human drivers. However, it is hard for policy networks to satisfy safety and feasibility constraints, which is not a challenging task for conventional motion generation methods, such as optimization-based approach. In this paper, we propose Constrained Policy Net (CPN), which can learn safe and feasible driving policy from arbitrary inequality-constrained optimization-based expert planners. Instead of supervised learning with L2norm as the loss, we incorporate the domain knowledge of the expert planner directly into the training loss of the policy net by applying barrier functions to the safety and feasibility constraints of the optimization problem. An exemplar scenario with obstacles on both sides is used to implement the proposed CPN. Test results demonstrate that the policy net can learn to generate motions near boundaries of safety and feasibility constraints to achieve high driving quality as the baseline optimization while the constraints are satisfied. Jiachen Li 0001, Yeping Hu, Masayoshi Tomizuka |
IECON | 4 |
| 2017 | Real-time robust finger gaits planning under object shape and dynamics uncertaintiesabstractDexterous manipulation has broad applications in assembly lines, warehouses and agriculture. To perform large-scale manipulation tasks for various objects, a multi-fingered robotic hand sometimes has to sequentially adjust its grasping gestures, i.e. the finger gaits, to address the workspace limits and guarantee the object stability. However, realizing finger gaits planning in dexterous manipulation is challenging due to the complicated grasp quality metrics, uncertainties on object shapes and dynamics (mass and moment of inertia), and unexpected slippage under uncertain contact dynamics. In this paper, a dual-stage optimization based planner is proposed to handle these challenges. In the first stage, a velocity-level finger gaits planner is introduced by combining object grasp quality with hand manipulability. The proposed finger gaits planner is computationally efficient and realizes finger gaiting without 3D model of the object. In the second stage, a robust manipulation controller using robust control and force optimization is proposed to address object dynamics uncertainties and external disturbances. The dual-stage planner is able to guarantee stability under unexpected slippage caused by uncertain contact dynamics. Moreover, it does not require velocity measurement or expensive 3D/6D tactile sensors. The proposed dual-stage optimization based planner is verified by simulations on Mujoco. The simulation video is available at [1]. Yongxiang Fan, Te Tang, Hsien-Chung Lin, Yu Zhao 0015, Masayoshi Tomizuka |
IROS | 5 |
| 2017 | State estimation for deformable objects by point registration and dynamic simulationabstractTo enhance the robotic manipulation of deformable objects, a robust state estimator is proposed to track the object configuration in real time. A Gaussian mixture model (GMM) is constructed to register the object nodes towards the noisy point cloud. To deal with occlusion, the coherent point drift (CPD) regularization is applied on the mixture model, so as to maintain the topological structure from previous sequences of data and to infer the object states in occluded area. The state estimation is further refined by running a dynamic simulation in parallel, which guarantees the estimates to satisfy the object's physical constraints. A series of rope tracking experiments are performed to evaluate the proposed state estimator. It is shown that the object can be tracked robustly with sensor noise, outliers and massive occlusion. Te Tang, Yongxiang Fan, Hsien-Chung Lin, Masayoshi Tomizuka |
IROS | 4 |
| 2017 | Boundary layer heuristic for search-based nonholonomic path planning in maze-like environmentsabstractAutomatic valet parking is widely viewed as a milestone towards fully autonomous driving. One of the key problems is nonholonomic path planning in maze-like environments (e.g. parking lots). To balance efficiency and passenger comfort, the planner needs to minimize the length of the path as well as the number of gear shifts. Lattice A* search is widely adopted for optimal path planning. However, existing heuristics do not evaluate the nonholonomic dynamic constraint and the collision avoidance constraint simultaneously, which may mislead the search. To efficiently search the environment, the boundary layer heuristic is proposed which puts large cost in the area that the vehicle must shift gear to escape. Such area is called the boundary layer. A simple and efficient geometric method to compute the boundary layer is proposed. The admissibility and consistency of the additive combination of the boundary layer heuristic and existing heuristics are proved in the paper. The simulation results verify that the introduction of the boundary layer heuristic improves the search performance by reducing the computation time by 56.1%. Changliu Liu, Yizhou Wang 0003, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 3 |
| 2017 | Speed profile planning in dynamic environments via temporal optimizationabstractTo generate safe and efficient trajectories for an automated vehicle in dynamic environments, a layered approach is usually considered, which separates path planning and speed profile planning. This paper is focused on speed profile planning for a given path that is represented by a set of waypoints. The speed profile will be generated using temporal optimization which optimizes the time stamps for all waypoints along the given path. The formulation of the problem under urban driving scenarios is discussed. To speed up the computation, the non-convex temporal optimization is approximated by a set of quadratic programs which are solved iteratively using the slack convex feasible set (SCFS) algorithm. The simulations in various urban driving scenarios validate the effectiveness of the method. Changliu Liu, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 3 |
| 2017 | Spatially-partitioned environmental representation and planning architecture for on-road autonomous drivingabstractConventional layered planning architecture temporally partitions the spatiotemporal motion planning by the path and speed, which is not suitable for lane change and overtaking scenarios with moving obstacles. In this paper, we propose to spatially partition the motion planning by longitudinal and lateral motions along the rough reference path in the Frenét Frame, which makes it possible to create linearized safety constraints for each layer in a variety of on-road driving scenarios. A generic environmental representation methodology is proposed with three topological elements and corresponding longitudinal constraints to compose all driving scenarios mentioned in this paper according to the overlap between the potential path of the autonomous vehicle and predicted path of other road users. Planners combining A* search and quadratic programming (QP) are designed to plan both rough long-term longitudinal motions and short-term trajectories to exploit the advantages of both search-based and optimization-based methods. Limits of vehicle kinematics and dynamics are considered in the planners to handle extreme cases. Simulation results show that the proposed framework can plan collision-free motions with high driving quality under complicated scenarios and emergency situations. Jianyu Chen 0002, Ching-Yao Chan, Changliu Liu, Masayoshi Tomizuka |
Intelligent Vehicles Symposium | 5 |
| 2016 | Algorithmic safety measures for intelligent industrial co-robotsabstractIn factories of the future, humans and robots are expected to be co-workers and co-inhabitants in the flexible production lines. It is important to ensure that humans and robots do not harm each other. This paper is concerned with functional issues to ensure safe and efficient interactions among human workers and the next generation intelligent industrial co-robots. The robot motion planning and control problem in a human involved environment is posed as a constrained optimal control problem. A modularized parallel controller structure is proposed to solve the problem online, which includes a baseline controller that ensures efficiency, and a safety controller that addresses real time safety by making a safe set invariant. Capsules are used to represent the complicated geometry of humans and robots. The design considerations of each module are discussed. Simulation studies which reproduce realistic scenarios are performed on a planar robot arm and a 6 DoF robot arm. The simulation results confirm the effectiveness of the method. Changliu Liu, Masayoshi Tomizuka |
ICRA | 2 |
| 2016 | Robust two-degree-of-freedom iterative learning control for flexibility compensation of industrial robot manipulatorsabstractMost industrial robots are actuated using geared motors with no direct load side measurement. The flexibility introduced by the gear reducer causes transmission errors and vibrations, which limits the adoption of robot manipulators in many demanding applications. This paper presents a lean and efficient scheme of iterative learning control (ILC) to compensate for the joint flexibility of industrial robot manipulators. A two-degree-of-freedom ILC method is introduced. Compared with the dual-stage ILC that has been previously proposed for servo flexibility compensation, the method is more effective and also enables a leaner implementation. In addition, in order to handle system variation, a robust synthesis method is developed by using H∞ and μ techniques in an innovative way. The proposed method is analyzed using simulation studies as well as tested on an actual industrial robot manipulator. Cong Wang 0015, Minghui Zheng, Masayoshi Tomizuka |
ICRA | 4 |
| 2016 | Robust impedance control with applications to a series-elastic actuated systemabstractImpedance control offers a theoretical basis for safe interaction between a robot and the environment, but model uncertainty, disturbances and actuation dynamics can compromise the accuracy of the rendered impedance in implementation. If both the interactive force and motion are directly sensed, the relationship between them can be robustly regulated to present the desired impedance dynamics. In this paper, a Disturbance Observer based controller architecture is presented which offers performance robustness for impedance control. Conditions for stability and passivity are developed, then this controller is analyzed on a series-elastic actuated system. The effect of actuation dynamics on both performance and stability is analyzed, then experimental results are presented. Kevin Haninger, Junkai Lu, Masayoshi Tomizuka |
IROS | 3 |
| 2016 | Human guidance programming on a 6-DoF robot with collision avoidanceabstractIn the application of physical human-robot interaction (pHRI), the collaboration between human and robot can significantly improve the production efficiency through combination of the human's flexible intelligence and the robot's consistent performance. In this application, however, it is an important concern to ensure the safety of the human and the robot. In the human guidance programming scenario, the operator plans a collision-free path for the robot end-effector, but the robot body might collide with an obstacle while being guided by the operator. In this paper, a novel on-line velocity based collision avoidance algorithm is developed to solve the problem in this particular scenario. The proposed algorithm gives an explicit solution to deal with both collision avoidance and human guidance command at the same time, which provides the operator a better and safer lead through programming experience. The real-time experiment is performed on FANUC LR Mate 200 iD/7L in three different obstacle scenarios. Hsien-Chung Lin, Yongxiang Fan, Te Tang, Masayoshi Tomizuka |
IROS | 4 |
| 2016 | Robotic manipulation of deformable objects by tangent space mapping and non-rigid registrationabstractRecent works of non-rigid registration have shown promising applications on tasks of deformable manipulation. Those approaches use thin plate spline-robust point matching (TPS-RPM) algorithm to regress a transformation function, which could generate a corresponding manipulation trajectory given a new pose/shape of the object. However, this method regards the object as a bunch of discrete and independent points. Structural information, such as shape and length, is lost during the transformation. This limitation makes the object's final shape to differ from training to test, and can sometimes cause damage to the object because of excessive stretching. To deal with these problems, this paper introduces a tangent space mapping (TSM) algorithm, which maps the deformable object in the tangent space instead of the Cartesian space to maintain structural information. The new algorithm is shown to be robust to the changes in the object's pose/shape, and the object's final shape is similar to that of training. It is also guaranteed not to overstretch the object during manipulation. A series of rope manipulation tests are performed to validate the effectiveness of the proposed algorithm. Te Tang, Changliu Liu, Masayoshi Tomizuka |
IROS | 4 |
| 2015 | Introduction and initial exploration of an Active/Passive Exoskeleton framework for portable assistanceabstractAssistive devices such as exoskeletons are capable of providing rehabilitative improvement and independence for individuals suffering from musculoskeletal conditions. Typical devices use either active assistance methods such as DC motors or passive methods such as springs. Active methods require a continuous power input, while passive methods are limited by user capability. This work introduces an Active/Passive EXoskeleton (APEX) framework. This device can passively provide continuous assistance, only requiring energy to change the dynamic properties of the passive state. The first prototype (APEX-α) is introduced and tested on six healthy subjects who performed hammer curls. It was found that changes in the passive state of the APEX-α affect the number of curls performed by an individual. By changing the passive state of the exoskeleton, increases in curl count of 65 - 92% were observed. This indicates the potential for such devices to provide assistance to an individual through the use of lightweight, energy efficient active/passive actuators. Robert Peter Matthew, Eric John Mica, Waiman Meinhold, Joel Alfredo Loeza, Masayoshi Tomizuka, Ruzena Bajcsy |
IROS | 5 |
| 2014 | Pose estimation in industrial machine vision systems under sensing dynamics: A statistical learning approachabstractThis paper deals with the problem of pose estimation (i.e., estimating position and orientation of an moving target) for real-time visual servoing, where the vision hardware is assumed to have severely limited measurement capability. In other words, we aim to compensate the slow sensor dynamics in industrial machine vision systems. The common approach is to predict the present target motion by propagating the delayed estimates with the target dynamics. Such method is sometimes problematic since the target motion characteristics (i.e., target dynamics) may change from one visual servoing task to another. Therefore, this paper presents a method which is able to estimate the target pose as well as learn the target dynamics. We apply the Expectation-Maximization algorithm to simultaneously solve the pose estimation problem and the target dynamics modeling problem. Several techniques including the extended Kalman filter/smoother, the block coordinate descent method, and the convex optimization method are utilized to address this problem. The effectiveness of the proposed algorithm is demonstrated experimentally on a 6-DOF industrial robot. Chung-Yen Lin, Cong Wang 0015, Masayoshi Tomizuka |
ICRA | 3 |
| 2014 | Design of kinematic controller for real-time vision guided robot manipulatorsabstractThis paper discusses the control strategies for robot manipulators to track moving targets based on realtime vision guidance. The work is motivated by some new demands of vision guided robot manipulators in which the workpieces being manipulated are in complex motion and the widely adopted look-then-move control strategy cannot give satisfactory performance. In order to serve industrial applications, the limited sensing and actuation capabilities have to be considered properly. In our work, a cascade control structure is introduced. The sensor and actuator limits are dealt with by consecutive modules in the controller respectively. In particular, the kinematic visual servoing (KVS) module is discussed in detail. It is a kinematic controller that generates reference trajectory in real-time. Sliding control is used to give a basic Jacobian-based design. Constrained optimal control is applied to address the actuator limits. validation is conducted through simulation and experiment. Cong Wang 0015, Chung-Yen Lin, Masayoshi Tomizuka |
ICRA | 3 |
| 2014 | Kinematic design and analysis for a macaque upper-limb exoskeleton with shoulder joint alignmentabstractAn exoskeleton design for a rhesus macaque subject is motivated, presented, and analyzed. As kinematic properties of the macaque's upper-limb have not been thoroughly studied, this paper introduces methods to determine properties relevant to exoskeleton design. Alignment with biological joints is critical for exoskeleton performance, but there are no accepted kinematic joint models for rhesus macaques. An algorithm is introduced which uses motion capture data to determine an appropriate model for the shoulder complex. An exoskeleton which incorporates this model is introduced, then analyzed. As joint speeds of macaques are also not well studied, a proposed analysis finds an upper bound on the joint speeds required to realize a given end effector speed in an arbitrary direction for all configurations within the workspace. Kevin Haninger, Junkai Lu, Masayoshi Tomizuka |
IROS | 4 |
| 2014 | Development of a rehabilitation robot suit with velocity and torque-based mechanical safety devicesabstractSafety is one of the most important issues in rehabilitation robot suits. We have proposed the structure of a rehabilitation robot suit equipped with two mechanical safety devices. The robot suit assists a patient's knee joint and the safety devices consist of only passive mechanical components without actuators, controllers, or batteries. We call one device the “velocity-based safety device” and the other the “torque-based safety device”. We expect a robot suit with the safety devices to be able to guarantee the safety even when the computer fails to operate functionally. In this paper, we begin by reviewing the characteristics of the safety devices and the structure of the rehabilitation robot suit equipped with the safety devices. Then we show a prototype robot suit developed based on the proposed structure. Finally, experimental results are demonstrated to verify the effectiveness of the safety devices in the prototype robot suit. Yoshihiro Kai, Satoshi Kitaguchi, Shotaro Kanno, Masayoshi Tomizuka |
IROS | 5 |
| 2014 | Modeling and controller design of cooperative robots in workspace sharing human-robot assembly teamsabstractHuman workers and robots are two major workforces in modern factories. For safety reasons, they are separated, which limits the productive potentials of both parties. It is promising if we can combine human's flexibility and robot's productivity in manufacturing. This paper investigates the modeling and controller design method of workspace sharing human-robot assembly teams and adopts a two-layer interaction model between the human and the robot. In theoretical analysis, enforcing invariance in a safe set guarantees safety. In implementation, an integrated method concerning online learning of closed loop human behavior and receding horizon control in the safe set is proposed. Simulation results in a 2D setup confirm the safety and efficiency of the algorithm. Masayoshi Tomizuka |
IROS | 2 |
| 2014 | Ensuring safety in human-robot coexistence environmentabstractThis paper proposes a safety index and an associated formulation in the optimization-based path planning framework to assess and ensure the safety of human workers in a human-robot coexistence environment. The safety index is evaluated using the ellipsoid coordinates (EC) attached to the robot links that represents the distance between the robot arm and the worker. To account for the inertial effect, the momentum of the robot links are projected onto the coordinates to generate additional measures of safety. The safety index is used as a constraint in the optimization problem so that a collision-free trajectory within a finite time horizon is generated online iteratively for the robot to move towards the desired position. To reduce the computational load for real-time implementation, the formulated optimization problem is further approximated by a quadratic problem. The safety index and the proposed formulations are simulated and validated in a two-link planar robot and the ITRI 7-DoF robot with a human worker moving inside the workspace of the robots. Chi-Shen Tsai, Jwu-Sheng Hu, Masayoshi Tomizuka |
IROS | 3 |
| 2014 | Fast planning of well conditioned trajectories for model learningabstractThis paper discusses the problem of planning well conditioned trajectories for learning a class of nonlinear models such as the imaging model of a camera and the multibody dynamic model of a robot. In such model learning problems, the model parameters can be linearly decoupled from system variables in the feature space. The learning accuracy and robustness against measurement noise and unmodeled response depend largely on the condition number of the data matrix. A new method is proposed to plan well conditioned trajectories efficiently by using low-discrepancy sequences and matrix subset selection. Application examples show promising results. Cong Wang 0015, Yu Zhao 0015, Chung-Yen Lin, Masayoshi Tomizuka |
IROS | 4 |
| 2014 | Improving Control Performance by Minimizing Jitter in RT-WiFi NetworksabstractWireless networked control systems have received significant attention due to their great advantages in enhanced system mobility, and reduced deployment and maintenance cost. To support a wide range of high-speed wireless control applications, we presented in our prior work the design and implementation of a flexible real-time high-speed wireless communication platform called RT-WiFi. RT-WiFi currently provides up to 6kHz sampling rate and deterministic timing guarantee on packet delivery. While guaranteed delivery latency is essential for networked control, control performance is also impacted by communication jitter and other QoS parameters. To reduce jitter, a flexible network manager is needed to control network-wide scheduling of packet transportation. In this paper, we present an RT-WiFi network manager design and propose efficient solutions for two fundamental RT-WiFi network management problems. To improve control performance in networked control systems, our RT-WiFi network manager is designed to generate data link layer communication schedule with minimum jitter under both static and dynamic network topologies. In order to minimize network management overhead, an efficient data structure called S-tree is invented to manage the communication requests to deal with network dynamics. We have implemented the RT-WiFi network manager, and validated its network and control performance through extensive experiments with a real application. Quan Leng, Yi-Hung Wei, Song Han 0002, Aloysius K. Mok, Masayoshi Tomizuka |
RTSS | 6 |
| 2013 | A nonlinear feedback controller for aerial self-righting by a tailed robotabstractIn this work, we propose a control scheme for attitude control of a falling, two link active tailed robot with only two degrees of freedom of actuation. We derive a simplified expression for the robot's angular momentum and invert this expression to solve for the shape velocities that drive the body's angular momentum to a desired value. By choosing a body angular velocity vector parallel to the axis of error rotation, the controller steers the robot towards its desired orientation. The proposed scheme is accomplished through feedback laws as opposed to feedforward trajectory generation, is fairly robust to model uncertainties, and is simple enough to implement on a miniature microcontroller. We verify our approach by implementing the controller on a small (175 g) robot platform, enabling rapid maneuvers approaching the spectacular capability of animals. Evan Chang-Siu, Thomas Libby, Robert J. Full, Masayoshi Tomizuka |
ICRA | 5 |
| 2013 | RT-WiFi: Real-Time High-Speed Communication Protocol for Wireless Cyber-Physical Control ApplicationsabstractApplying wireless technologies in control systems can significantly enhance the system mobility and reduce the deployment and maintenance cost. Existing wireless technology standards, however either cannot provide real-time guarantee on packet delivery or are not fast enough to support high-speed control systems which typically require 1kHz or higher sampling rate. Nondeterministic packet transmission and insufficiently high sampling rate will severely hurt the control performance. To address this problem, in this paper, we present our design and implementation of a real-time high-speed wireless communication protocol called RT-WiFi. RT-WiFi is a TDMA data link layer protocol based on IEEE 802.11 physical layer to provide deterministic timing guarantee on packet delivery and high sampling rate up to 6kHz. It incorporates configurable components for adjusting design trade-offs including sampling rate, latency variance, reliability, and compatibility to existing Wi-Fi networks, thus can serve as an ideal communication platform for supporting a wide range of high-speed wireless control systems. We implemented RT-WiFi on commercial off-the-shelf hardware and integrated it into a mobile gait rehabilitation system. Our extensive experiments demonstrate the effectiveness of RT-WiFi in providing deterministic packet delivery in both data link layer and application layer, which further eases the controller design and significantly improve the control performance. Yi-Hung Wei, Quan Leng, Song Han 0002, Aloysius K. Mok, Masayoshi Tomizuka |
RTSS | 6 |
| 2012 | Compensation of packet loss for a network-based rehabilitation systemabstractIn this paper, a network-based rehabilitation system is proposed to increase mobility of a rehabilitation system and to enable tele-rehabilitation. Control algorithms and rehabilitation strategies distributed at the central location (physical therapist) and the local site (patient) communicate over wireless network to realize a network-based rehabilitation system. To deal with possible packet losses over wireless network, a modified linear quadratic Gaussian (LQG) controller and a disturbance observer (DOB) are applied. The performance of the proposed system and control algorithms is verified by simulation and experiment with an actual knee rehabilitation system. The simulation and experiment results show that the network-based rehabilitation system with the proposed control schemes can generate the desired assistive torque accurately in presence of packet losses. Joonbum Bae, Masayoshi Tomizuka |
ICRA | 3 |
| 2012 | A sensor-based approach for error compensation of industrial robotic workcellsabstractIndustrial robotic manipulators have excellent repeatability while accuracy is significantly poorer. Numerous error sources in the robotic workcell contributes to the accuracy problem. Modeling and identification of all the errors to achieve the required levels of accuracy may be difficult. To resolve the accuracy issues, a sensor based indirect error compensation approach is proposed in this paper where the errors are compensated online via measurements of the work object. The sensor captures a point cloud of the work object and with the CAD model of the work object, the actual relative pose of the sensor frame and work object frame can be established via a point cloud registration. Once this relationship has been established, the robot will be able to move the tool accurately relative to the work object frame near the point of compensation. A data pre-processing technique is proposed to reduce computation time and prevent a local minima solution during point cloud registration. A simulation study is presented to illustrate the effectiveness of the proposed solution. Pey Yuen Tao, Guilin Yang, Masayoshi Tomizuka |
ICRA | 3 |
| 2012 | Robot end-effector sensing with position sensitive detector and inertial sensorsabstractFor the motion control of industrial robots, the end-effector performance is of the ultimate interest. However, industrial robots are generally only equipped with motor-side encoders. Accurate estimation of the end-effector position and velocity is thus difficult due to complex joint dynamics. To overcome this problem, this paper presents an optical sensor based on position sensitive detector (PSD), referred as PSD camera, for direct end-effector position sensing. PSD features high precision and fast response while being cost-effective, thus is favorable for real-time feedback applications. In addition, to acquire good velocity estimation, a kinematic Kalman filter (KKF) is applied to fuse the measurement from the PSD camera with that from inertial sensors mounted on the end-effector. The performance of the developed PSD camera and the application of the KKF sensor fusion scheme have been validated through experiments on an industrial robot. Cong Wang 0015, Masayoshi Tomizuka |
ICRA | 3 |
| 2011 | A lizard-inspired active tail enables rapid maneuvers and dynamic stabilization in a terrestrial robotabstractWe present a novel approach to stabilizing rapid locomotion in mobile terrestrial robots inspired by the tail function of lizards.We built a 177 (g) robot with inertial sensors and a single degree-of-freedom active tail. By utilizing both contact forces and zero net angular momentum maneuvering, our tailed robot can rapidly right itself in a fall, avoid flipping over after a large perturbation, and smoothly transition between surfaces of different slopes. We also use a modeling approach to show that a tail-like design offers significant advantages to other alternatives, including reaction wheels, when the speed of response is important. Evan Chang-Siu, Thomas Libby, Masayoshi Tomizuka, Robert J. Full |
IROS | 3 |
| 2011 | Time-varying complementary filtering for attitude estimationabstractComplementary filtering (CF) is a well known method that can effectively fuse a gyroscope and accelerometer measurement in order to robustly estimate the attitude of a rigid body in a planar single degree of freedom (DOF) setting. The attitude can be estimated individually by either integrating the gyroscope measurement or by calculating the inverse tangent of the components of a 2-axis accelerometer. The gyroscope can adequately estimate the angle in the higher frequency region, but suffers from drift issues at low frequency, whereas the accelerometer can accurately measure the acceleration and thus direction of gravity, but loses this accuracy when faced with motion accelerations. CF traditionally uses linear time invariant filters, however, this paper presents an extension to the CF method by proposing time-varying parameters. A fuzzy logic method is developed to adjust the parameters. Stability analysis as well as experimental results are presented to verify the proposed method. Evan Chang-Siu, Masayoshi Tomizuka, Kyoungchul Kong |
IROS | 2 |
| 2010 | A compact rotary series elastic actuator for knee joint assistive systemabstractPrecise and large torque generation, back-drivability, low output impedance, and compactness of hardware are important requirements for human assistive robots. In this paper, a compact rotary series elastic actuator (cRSEA) is designed considering these requirements. To magnify the torque generated by an electric motor in the limited space of the compact device, a worm gear is utilized. However, the actual torque amplification ratio provided by the worm gear is different from the nominal speed reduction ratio due to friction, which makes the controller design challenging. In this paper, the friction effect is considered in the model of cRSEA, and a robust control algorithm is designed to precisely control the torque output in the presence of nonlinearities such as the friction. The mechanical design and dynamic model of the proposed device and the design of a robust control algorithm are discussed, and actuation performance is verified by experiments. Kyoungchul Kong, Joonbum Bae, Masayoshi Tomizuka |
ICRA | 3 |
| 2010 | Fuzzy Stabilization of Nonlinear Systems under Sampled-Data Feedback: An Exact Discrete-Time Model ApproachabstractThis paper addresses stabilization problems for a nonlinear system via a sampled-data fuzzy controller. The nonlinear system is assumed to be exactly modeled in Takagi-Sugeno's form, at least locally. Unlike the conventional direct discrete-time design approach based on an approximate discrete-time model, the sampled-data fuzzy controllers are designed based on an exact discrete-time model. Sufficient design conditions are formulated in terms of linear matrix inequalities. It is shown that whenever the exact discrete-time fuzzy model is asymptotically stabilizable via the sampled-data fuzzy controller uniformly bounded in the state, then so is the original nonlinear system. A numerical example is given to illustrate the effectiveness of the proposed methodology. Do Wan Kim, Ho Jae Lee, Masayoshi Tomizuka |
IEEE Trans. Fuzzy Syst. | 3 |
| 2009 | Design of a rehabilitation device based on a mechanical link systemabstractRealizing an ideal impedance control system in lower extremity rehabilitation systems is challenged by mechanical impedance of robot hardware. Although some studies in the field of control systems have been helpful in reducing the mechanical impedance of actuators, they have not been able to remove the inertia of robot hardware. This paper introduces an alternative design in which mechanical links are utilized. The mechanical links are driven by one actuator without any complicated servo systems. The design parameters are optimized for the link system to generate the normal walking motion. The simulation data shows that the normal gait patterns are realized successfully. The device is connected to a patient using elastic components, and therefore the inertia of the robot is not directly imposed on the patient. The patient's legs are guided to follow the motion of the robot with the forces generated by the elastic components. Kyoungchul Kong, Chulhyun Baek, Masayoshi Tomizuka |
ICRA | 3 |
| 2009 | Robotic rehabilitation treatments: Realization of aquatic therapy effects in exoskeleton systemsabstractExoskeletons are attracting a great attention as a new means of rehabilitation devices. In such applications, control algorithms of exoskeletons are often inspired by nature for natural and effective assistance for patients. In this paper, a control algorithm is inspired by aquatic therapy. Aquatic therapy has various benefits for rehabilitation processes based on useful properties of water, e.g. buoyancy and drag. However, realization of such effects is challenged by limitations in hardware, such as mechanical impedance or impreciseness of actuator forces. Therefore, the resistive forces generated by actuators, which cause serious discomfort to patients, are precisely modeled and compensated to realize the control algorithm inspired by aquatic therapy effectively. The proposed methods are implemented in SUBAR developed by Sogang University and verified by experiments. Kyoungchul Kong, Hyosang Moon, Beomsoo Hwang, Doyoung Jeon, Masayoshi Tomizuka |
ICRA | 5 |
| 2009 | Impedance Compensation of SUBAR for Back-Drivable Force-Mode ActuationabstractThe Sogang University biomedical assistive robot (SUBAR), which is an advanced version of the exoskeleton for patients and the old by Songang (EXPOS) is a wearable robot developed to assist physically impaired people. It provides a person with assistive forces controlled by human intentions. If a standard geared DC motor is applied, however, the control efforts will be used mainly to overcome the resistive forces caused by the friction, the damping, and the inertia in actuators. In this paper, such undesired properties are rejected by applying a flexible transmission. With the proposed method, it is intended that an actuator exhibits zero impedance without friction while generating the desired torques precisely. Since the actuation system of SUBAR has a large model variation due to human-robot interaction, a control algorithm for the flexible transmission is designed based on a robust control method. In this paper, the mechanical design of SUBAR, including the flexible transmission and its associated control algorithm, are presented. They are also verified by experiments. Kyoungchul Kong, Hyosang Moon, Beomsoo Hwang, Doyoung Jeon, Masayoshi Tomizuka |
IEEE Trans. Robotics | 5 |
| 2008 | Smooth and continuous human gait phase detection based on foot pressure patternsabstractMeasurement of ground contact forces (GCF) provides necessary information to detect human gait phases. In this paper, a new analysis method of the GCF signals is discussed for detection of the gait phases. Human gaits are complicated, and the gait phases can not be exactly distinguished by comparing sensor outputs to a threshold. This paper mainly discusses how to detect the gait phases continuously and smoothly. The proposed analysis method is intended for applications to power assistive devices for patients, as well as diagnostics of pathological gait. Smooth and continuous detection of the gait phases enables a full use of information obtained from GCF sensors. For experimental verification, smart shoes have been developed. Each smart shoe has four GCF sensors embedded between the cushion pad and the sole. The performances are experimentally verified for both normal and abnormal gaits, and a means for quantification of abnormalities in the gait is also introduced in this paper. Kyoungchul Kong, Masayoshi Tomizuka |
ICRA | 2 |
| 2007 | Flexible Joint Actuator for Patient's Rehabilitation DeviceabstractRehabilitation devices require a very precise actuating system. In this paper, a flexible joint actuator is proposed as an actuating system of an intelligent active orthosis. To generate joint torque as desired, a spring is installed between a motor and human joint and the motor is controlled to have a proper spring deflection for torque control. When the desired torque is zero, the motor should follow human joint motion which requires that the friction and inertia of the motor are compensated. The human joint and body part represent the load to the flexible joint actuator. They interact with environment and their parameters are not fixed. The controller for the flexible joint actuator must operate under these conditions. Kyoungchul Kong, Masayoshi Tomizuka |
RO-MAN | 2 |
| 2007 | Friction modelling and compensation for motion control using hybrid neural network models
M. Kemal Ciliz, Masayoshi Tomizuka |
Eng. Appl. Artif. Intell. | 2 |
| 2002 | A navigation system for unmanned vehicles in automated highway systemsabstractIn this paper, a new method for generating a set of collision-free maneuvers for unmanned vehicles in automated highway systems is introduced. The low computational cost of the proposed maneuver planner allows its execution as frequent as the new positions and speeds of vehicles (obstacles) in sight are received. This planner is based on the computation of the minimum translational distance between two mobile objects. This distance is then used for predicting and avoiding a collision, generating (if it is possible and safe), overtaking (changing to the left and right lane), braking and keeping-the-lane maneuvers. These maneuvers can be provided to a decision-making process for selecting the most suitable one. Enrique J. Bernabeu, Josep Tornero, Masayoshi Tomizuka |
IROS | 3 |
| 2001 | Fuzzy Logic Modeling & Control for Drilling Composite LaminatesabstractIn drilling of composite laminates, it is important to minimize or reduce occurrences of delaminations. In particular, a peel-up delamination at entrance and push-out delamination at exit are common. Delaminations may be avoided by regulating the drill thrust force which can be controlled by adjusting the feedrate of the drill. Dynamics involved in drilling of composite laminates is time varying and nonlinear. In this paper, a fuzzy logic model and control strategy are proposed. Simulation results show that the fuzzy model can describe the nonlinear time-varying process well. The fuzzy controller realizes a fast rise time and a little overshoot of drilling force. Byeong-Mook Chung, Masayoshi Tomizuka |
FUZZ-IEEE | 2 |
| 2001 | Collision Prediction and Avoidance Amidst Moving Objects for Trajectory Planning ApplicationsabstractA methodology for computing a collision-free trajectory for mobile robots amidst moving objects is presented. This planner is based on a technique for computing the minimum translational distance between two mobile objects. This distance is then used for predicting and avoiding a collision. The computation of this distance is based on the application of the GJK algorithm to a particular subset of the Minkowski difference set of the involved objects. This subset states the separation or penetration distance between two objects along their given motions. When a collision is predicted, a collision-free intermediate temporal-position is generated avoiding such a collision. Enrique J. Bernabeu, Josep Tornero, Masayoshi Tomizuka |
ICRA | 3 |
| 2000 | Robust adaptive control using a universal approximator for SISO nonlinear systemsabstractThis paper deals with the robust adaptive control of a class of nonlinear systems in the presence of parametric uncertainties and dominant uncertain nonlinearities. The proposed controller utilizes the robust adaptive control to guarantee uniform boundedness and convergence of tracking errors. In addition, an adaptive fuzzy logic system is used as a universal approximator to reduce the model uncertainties coming from uncertain nonlinearities and to improve tracking performance. The approach does not require the matching condition imposed on control systems by using the backstepping design procedure, and provides boundedness of tracking errors under poor parameter adaptation. The method can be applied to a class of single-input single-output (SISO) nonlinear systems, transformable to a parametric-strict-feedback form. Hyeongcheol Lee, Masayoshi Tomizuka |
IEEE Trans. Fuzzy Syst. | 2 |
| 1995 | Adaptive Control of Robot Manipulators with Anti-Backlash GearsabstractThis paper proposes a control method for robot manipulators that have anti-backlash gears in the joints. The anti-backlash gear is modeled as a three segment flexible joint characteristic. A controller consists of a PD feedback part and a full dynamics feedforward part. This control method employs the sliding control scheme. In order to make the tip of the manipulator track desired trajectories and to eliminate steady state errors, this method utilizes two sliding surfaces for both links and actuators. For this purpose, positions and velocities of both links and actuators are fedback. An anti-backlash inverse characteristic is introduced to obtain the desired trajectories for the actuators from the desired trajectories of the links. An adaptation scheme to the variation of payloads is also investigated. Naoki Imasaki, Masayoshi Tomizuka |
ICRA | 2 |
| 1995 | Adaptive Control of Two Robot Arms Carrying an Unknown ObjectabstractWe deal with the control of two robot manipulators carrying an unknown object. We consider the uncertainties in mass, inertia and center of mass of the object. A dynamic model is obtained for the combined system. An adaptive scheme is proposed to track the desired motion trajectory of the object in the presence of uncertainties. A feedforward plus PI type force controller is designed for internal force control. A Lyapunov based stability analysis is conducted for the proposed scheme. In the adaptive motion control scheme for the combined system we cannot specify the individual torques of the manipulators uniquely. To obtain a set of individual torques an optimal control problem is formulated by minimizing a cost function. Prabhakar R. Regilla, Masayoshi Tomizuka |
ICRA | 2 |
| 1995 | Robust Adaptive Constrained Motion and Force Control of Manipulators with Guaranteed Transient PerformanceabstractIn this paper, a combined adaptive and smooth sliding mode control (SMC) methodology is suggested to design a constrained motion and force tracking controller so as to achieve asymptotic motion and force tracking without persistent excitation condition in the presence of parametric uncertainties. Filtered error signals are used in both motion and force sliding surfaces to enhance dynamic response of the system. The proposed controller provides guaranteed transient performance with prescribed final tracking accuracy in the presence of both parametric uncertainties and external disturbances or modelling errors. Simulation results are presented to illustrate the proposed controller. Masayoshi Tomizuka |
ICRA | 2 |
| 1994 | Robust Desired Compensation Adaptive Control of Robot Manipulators with Guaranteed Transient PerformanceabstractA robust adaptive controller with guaranteed transient performance under a desired compensation adaptation law is developed for trajectory tracking control of robot manipulators in the presence of parametric uncertainties and external disturbances. With some modifications to the conventional adaptation law, a new control law is redesigned by combining the design methodologies of adaptive control and continuous sliding mode control. The suggested controller preserves the advantages of the true methods, namely, asymptotic stability of adaptive system for parametric uncertainties and globally uniformly ultimately bounded stability with guaranteed transient performance of sliding mode control for both parametric uncertainties and external disturbances. The control law is continuous, and the chattering problem of sliding mode control is avoided. Experimental results illustrate the effectiveness of the proposed methods.> Masayoshi Tomizuka |
ICRA | 2 |
| 1994 | Fuzzy smoothing algorithms for variable structure systemsabstractA variable structure system (VSS) is a control system implementing different control laws in different regions of the state space divided by a set of boundary manifolds. The control input switches from one control law to another when the state crosses the boundary manifolds. In general, the control input may not be smooth when switching at these boundary manifolds and may excite high frequency dynamics. This paper proposes two fuzzy rule based algorithms for smoothing the control input. The merits of these fuzzy smoothing control algorithms are illustrated by two examples: a semiactive suspension system based on optimal control and a direct drive robot arm under discrete time sliding mode control. The controller design for these two examples is a blend of traditional control theoretic approaches and fuzzy rule based approaches.> Yean-Ren Hwang, Masayoshi Tomizuka |
IEEE Trans. Fuzzy Syst. | 2 |
| 1993 | Learning hybrid force and position control of robot manipulatorsabstractThe learning control is applied to hybrid force and position control of robot manipulators. When the geometry and position of a constraint surface is known, the hybrid force and position controller and the feedforward compensator can be designed in the constraint coordinates. When the operation is periodic, the learning hybrid force and position control enhance the control performance as the feedforward compensator is updated in each cycle by the force and position error in the preceding trials. This scheme is proved to be asymptotically stable. A two degree of freedom SCARA-type direct-drive robot manipulator is used to test the learning hybrid force and position control. The deburring tool mounted on the upper link of the robot could follow a flat, tilted flat, and curved 1/4 " aluminum plate with a desired contact force of 10 N (within the root-mean-square force error of 1.95 N) and with a desired tangential velocity. The experiments confirmed the effectiveness of the learning hybrid force and position controller.> Doyoung Jeon, Masayoshi Tomizuka |
IEEE Trans. Robotics Autom. | 2 |
| 1993 | Fuzzy gain scheduling of PID controllersabstractThis paper describes the development of a fuzzy gain scheduling scheme of PID controllers for process control. Fuzzy rules and reasoning are utilized online to determine the controller parameters based on the error signal and its first difference. Simulation results demonstrate that better control performance can be achieved in comparison with Ziegler-Nichols controllers and Kitamori's PID controllers.> Zhen-Yu Zhao, Masayoshi Tomizuka, Satoru Isaka |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1992 | Learning hybrid force and position control of robot manipulatorsabstractA hybrid force and position controller with a feedforward compensator is designed in the constraint coordinates. When the operation is periodic, the learning control is applied to enhance the performance of the hybrid force and position control as the feedforward compensator is updated in each cycle by the force and position error in the preceding trials. This scheme is proved to be stable. In the experiments, a two-degree-of-freedom SCARA-type direct-drive robot manipulator was used. The deburring tool mounted on the upper link of the robot could follow a flat, tilted flat, and curved 1/4 inch aluminum plate with a desired contact force of 10 N, and with desired tangential velocity. Considering the loss of contact observed at the initial trial, the performance of the system improved significantly.> Doyoung Jeon, Masayoshi Tomizuka |
ICRA | 2 |
| 1990 | Plug in repetitive control for industrial robotic manipulatorsabstractThe implementation of plug-in repetitive control on the direct drive axis of a prototype GMF A500 robot is considered. Plug-in repetitive control was developed and implemented using an IBM AT. A multirate sampling algorithm was designed to reduce the memory requirements of repetitive control. A microcontroller board based on the Intel 8096 was designed so that plug-in repetitive control could be applied from a small unit embedded within the KAREL controller. In each application, experimental results show that the tracking error is smoothly absorbed in a few cycles.> C. Cosner, George Anwar, Masayoshi Tomizuka |
ICRA | 3 |
| 1990 | Trajectory planning for coordinated motion of a robot and a positioning table. II. Optimal trajectory specificationabstractFor pt.I see ibid., p.735-45 (1990). A robot and positioning table system is a kinematically redundant system with respect to planar motion. Two strategies were developed in pt.I to resolve this redundancy and to specify path shapes that make the best utilization of the workspace and speed characteristics of the two devices. In this part, a one-variable dynamic programming approach is developed to obtain a near-minimum time and/or energy trajectory of the two devices. The developed strategies were studied using this path-planning algorithm on a model of a 3 d.o.f. robot and a two-axis linear positioning table for a variety of path shapes and constraints. In addition, experiments were performed in a typical workcell to show the feasibility of the developed strategies. In moving the two devices in opposite directions, the least travel time is obtained when the original path is resolved more in favor of the faster device, whereas the quality of resultant motion is dependent on the controller performance of each device.> Musa K. Jouaneh, David A. Dornfeld, Masayoshi Tomizuka |
IEEE Trans. Robotics Autom. | 3 |
| 1990 | A self-paced fuzzy tracking controller for two-dimensional motion controlabstractA heuristic controller is presented that takes the form of a set of fuzzy linguistic rules. Simulations show that the fuzzy logic controller (FLC) yields better results than the conventional PD controller. A self-paced fuzzy tracking controller (SPFTC) designed for two-dimensional path tracking is also presented. The SPFTC adjusts the tracking speed in accordance with contour conditions such as curvature; the fuzzy-rule-based adjustment of the tracking speed improves performance in terms of tracking precision and travel time. Advantages of the FLC and SPFTC are demonstrated by a simulation study.> Liang-Jong Huang, Masayoshi Tomizuka |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1989 | Model reference adaptive control and repetitive control for robot manipulatorsabstractA combination of adaptive and repetitive control for dynamic control of robot manipulators is introduced. The repetitive controller is a plug-in controller for connection to an existing model-reference adaptive controller. The input to the repetitive controller is the same as the adaptation error signal for the adaptive controller. In this scheme, desired trajectories and external disturbance are assumed to be periodic. Global stability for this control system is achieved and disturbance is rejected successfully. Based on the characteristics of the reference model, two types of control algorithm are presented: a unity-gain and a pure-integrator reference model. Discretized versions suitable for digital implementations, and sufficient conditions for their stability in the discrete time domain are presented. The performance of the adaptive and repetitive controller is demonstrated by simulation.> Ming-Chang Tsai, Masayoshi Tomizuka |
ICRA | 2 |
| 1988 | Discrete time repetitive control for robot manipulatorsabstractThe analysis and experimental implementation of a discrete-time repetitive control scheme is presented. The repetitive control structure is designed to be easily implemented on any system without modification to the existing controller. Simulation and experimental results show that the repetitive controller in conjunction with a computed-torque control or a simple proportional-derivative control law achieves good tracking performance when the desired trajectory is periodic and the period is known.> Ming-Chang Tsai, George Anwar, Masayoshi Tomizuka |
ICRA | 3 |
| 1987 | Model reference adaptive control of a two axis direct drive manipulator armabstractThe model reference adaptive control of a two axis direct drive manipulator arm is presented. The two axes exhibit a significant dynamic interaction. The model reference adaptive controller is utilized in the velocity loop to adaptively decouple the dynamic interaction and to linearize the dynamics. The outer loop controller for positioning and tracking is designed based on the linear decoupled dynamics. The adaptive velocity loop controller as well as position loop controller are implemented digitally. The use of a series-parallel, as well as a parallel reference model is suggested in the adaptive velocity loop controller. The stability problems arising from the straightforward digital implementation of the continuous time algorithm are analyzed and a modification in the digital algorithm is introduced which guarantees the asymptotic stability of the system. Simulation and experimental results show a consistently superior performance of the manipulator under adaptive control. Roberto Horowitz, Ming-Chang Tsai, George Anwar, Masayoshi Tomizuka |
ICRA | 4 |
| 1987 | Control of tool/workpiece contact force with application to robotic deburringabstractThe design and implementation of a microprocessor-based system to control the interaction forces between a five-axis articulated robot and a workpiece is described. The control system worked in parallel with a robot controller by calculating position corrections that allowed forces to be controlled in the desired manner. These corrections were successfully interfaced to the controller's position control loop on an individual-axis level. Stable force-control algorithms were designed in spite of limitations imposed by flexibility in the robot drive train. For multi-degree-of-freedom force control, it is shown that each axis can be considered autonomous, obviating the need for a multivariable approach. Force control was implemented in both edge following and deburring experiments. In edge following, the commanded normal force ranged from 1 to 15 N, while the root mean square (rms) force errors remained constant. Errors increased from 0.5 to 1.5 N rms as tangential speed was increased from 1 to 9 cm s-1. The performance of the force control system during deburring operations was characterized across the full force and speed range of the cutting tools used. The smoothness of cut was shown to be consistent with manual deburring operations in terms of optimal feed and metal removal rates. Tomasz Stepien, Larry Sweet, Malcolm C. Good, Masayoshi Tomizuka |
IEEE J. Robotics Autom. | 4 |
| 1986 | Application of nonlinear friction compensation to robot arm controlabstractThis paper deals with a design of a digital servo controller for a robot manipulator with mechanical nonlinearities. The experimental system is a D.C. motor driven manipulator arm. A model to predict the nonlinear behavior exhibited by the experimental system and a nonlinear controller with Coulomb friction compensator are developed and tested. The effectiveness of the nonlinear controller is certified by designing a linear tracking controller based on this scheme. Experimental and simulated responses are given. Tomoaki Kubo, George Anwar, Masayoshi Tomizuka |
ICRA | 3 |
| 1985 | Control of tool/Workpiece contact force with application to robotic deburringabstractThe design and implementation of a microprocessor based control system to control the interaction forces between a five axis articulated robot and a workpiece is described. The control system works in parallel with a robot controller by calculating position corrections that allow force to be controlled in the desired manner. The corrections were successfully interfaced to the position control loop on an individual axis level. Stable force control algorithms were designed in spite of limitations imposed by flexibility in the robot drive train. In the multi-degree of freedom control case it is shown that each axis can be considered autonomous, obviating the need for a multivariable approach. Control was implemented in edge following experiments. Across levels of commanded normal force ranging from 0 to 15 N, the RMS force errors remain constant. Errors increased from 0.5 N to 1.5 N RMS as tangential speed was increased from 0 to 9 cm/sec. The performance of the force control system during deburring operations is characterized across the full force and speed range of the cutting tools used. Smoothness of cut is shown to be consistent with standard deburring operations in terms of optimal feed and metal removal rates. Tomasz Stepien, Larry Sweet, Malcolm C. Good, Masayoshi Tomizuka |
ICRA | 4 |