Behzad Dariush

dblp:10/1984 · DBLP profile ↗
← Back
35ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 8 first-author · 13 since 2021Systems, architecture and hardware · 17 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Towards Driver Behavior Understanding: Weakly-Supervised Risk Perception in Driving Scenes
Nakul Agarwal, Yi-Ting Chen 0001, Behzad Dariush
IV3
2025 Task-Aware Resolution Optimization for Visual Large Language Models
abstract
Real-world vision-language applications demand varying levels of perceptual granularity.However, most existing visual large language models (VLLMs), such as LLaVA, preassume a fixed resolution for downstream tasks, which leads to subpar performance.To address this problem, we first conduct a comprehensive and pioneering investigation into the resolution preferences of different visionlanguage tasks, revealing a correlation between resolution preferences with ❶ image complexity, and ❷ uncertainty variance of the VLLM at different image input resolutions.Building on this insight, we propose an empirical formula to determine the optimal resolution for a given vision-language task, combining these two factors.Second, based on rigorous experiments, we propose a novel parameter-efficient fine-tuning technique to extend the visual input resolution of pre-trained VLLMs to the identified optimal resolution.Extensive experiments on various vision-language tasks validate the effectiveness of our method.
Weiqing Luo, Zhen Tan 0001, Kwonjoon Lee, Behzad Dariush, Tianlong Chen 0001
EMNLP6
2025 COMBO: Compositional World Models for Embodied Multi-Agent Cooperation
abstract
In this paper, we investigate the problem of embodied multi-agent cooperation, where decentralized agents must cooperate given only egocentric views of the world. To effectively plan in this setting, in contrast to learning world dynamics in a single-agent scenario, we must simulate world dynamics conditioned on an arbitrary number of agents' actions given only partial egocentric visual observations of the world. To address this issue of partial observability, we first train generative models to estimate the overall world state given partial egocentric observations. To enable accurate simulation of multiple sets of actions on this world state, we then propose to learn a compositional world model for multi-agent cooperation by factorizing the naturally composable joint actions of multiple agents and compositionally generating the video conditioned on the world state. By leveraging this compositional world model, in combination with Vision Language Models to infer the actions of other agents, we can use a tree search procedure to integrate these modules and facilitate online cooperative planning. We evaluate our methods on three challenging benchmarks with 2-4 agents. The results show our compositional world model is effective and the framework enables the embodied agents to cooperate efficiently with different agents across various tasks and an arbitrary number of agents, showing the promising future of our proposed methods. More videos can be found at https://umass-embodied-agi.github.io/COMBO
Qiushi Lyu, Sunli Chen, Tianmin Shu, Behzad Dariush, Kwonjoon Lee, Yilun Du, Chuang Gan 0001
ICLR7
2025 Generalized Mission Planning for Heterogeneous Multi-Robot Teams via LLM-Constructed Hierarchical Trees
abstract
We present a novel mission-planning strategy for heterogeneous multi-robot teams, taking into account the specific constraints and capabilities of each robot. Our approach employs hierarchical trees to systematically break down complex missions into manageable sub-tasks. We develop specialized APIs and tools, which are utilized by Large Language Models (LLMs) to efficiently construct these hierarchical trees. Once the hierarchical tree is generated, it is further decomposed to create optimized schedules for each robot, ensuring adherence to their individual constraints and capabilities. We demonstrate the effectiveness of our framework through detailed examples covering a wide range of missions, showcasing its flexibility and scalability.
David Isele, Enna Sachdeva, Pin-Hao Huang, Behzad Dariush, Kwonjoon Lee, Sangjae Bae
ICRA5
2025 Edit Distance Based Intention Estimation for Teleoperated Assembly
abstract
We address the problem of intention estimation in human-robot teleoperation, which involves identifying the task being completed and predicting the next actions. Our approach sequentially quantifies the similarity between the observed action sequence and nominal action sequences representing possible tasks using the edit distance metric. Task estimation and action prediction are then performed using a nearest-neighbor rule. A key advantage of our approach is its robustness to deviations in operator actions and action recognition errors, commonly encountered in real-world teleoperation settings. Through extensive experiments on both real and simulated data, we demonstrate that our method largely outperforms alternative approaches, including probabilistic graphical models and transformer-based methods, particularly in scenarios with significant action deviations or action recognition errors. Additionally, we construct task distance matrices to analyze task similarities and potential confusion points, providing insights into when and where estimation errors are likely to occur. This analysis can guide the design of more distinctive task sequences and further improve the reliability of teleoperated robotic systems.
Aolin Xu 0002, Songpo Li, Prakash Baskaran, Soshi Iba, Behzad Dariush
IROS5
2025 A Probabilistic Programming Approach to Intention Estimation in Human-Robot Teleoperated Assembly Tasks
abstract
We propose a new approach to solving the problem of intention estimation in human-robot teleoperation for assembly tasks, which includes task estimation and action prediction. Our approach uses probabilistic graphical models to represent the joint distribution of the task and the actions to be taken to complete the task. Both model learning and inference are implemented with Pyro, a state-of-the-art probabilistic programming language. The distinctive feature from the traditional hidden Markov model type of probabilistic methods is that our model takes the time information into account and explicitly models the individual distributions of all the variables under consideration. By doing this, we fully utilize the power of probabilistic programming, and achieve accurate distribution hence uncertainty estimations. Working with a pretrained action recognition module, the proposed model can be trained solely on a tiny instruction manual of the assembly tasks and can be retrained with minimal overhead whenever the manual is changed or augmented, avoiding the need for the costly data reannotation and retraining by the end-to-end learning based methods. We also compare our method with a transformer based model trained directly on the instruction manual, and our method shows superior accuracy in both intention estimation and their distribution estimations. We additionally identify failure cases of both our method and the transformer-based method, and envision methods for improvement.
Aolin Xu 0002, Songpo Li, Prakash Baskaran, Karankumar Patel, Soshi Iba, Behzad Dariush
IROS6
2024 Follow the Rules: Reasoning for Video Anomaly Detection with Large Language Models
Yuchen Yang 0001, Kwonjoon Lee, Behzad Dariush, Yinzhi Cao, Shao-Yuan Lo
ECCV (81)3
2024 Optimal Driver Warning Generation in Dynamic Driving Environment
abstract
The driver warning system that alerts the human driver about potential risks during driving is a key feature of an advanced driver assistance system. Existing driver warning technologies, mainly the forward collision warning and unsafe lane change warning, can reduce the risk of collision caused by human errors. However, the current design methods have several major limitations. Firstly, the warnings are mainly generated in a one-shot manner without modeling the ego driver’s reactions and surrounding objects, which reduces the flexibility and generality of the system over different scenarios. Additionally, the triggering conditions of warning are mostly rule-based threshold-checking given the current state, which lacks the prediction of the potential risk in a sufficiently long future horizon. In this work, we study the problem of optimally generating driver warnings by considering the interactions among the generated warning, the driver behavior, and the states of ego and surrounding vehicles on a long horizon. The warning generation problem is formulated as a partially observed Markov decision process (POMDP). An optimal warning generation framework is proposed as a solution to the proposed POMDP. The simulation experiments demonstrate the superiority of the proposed solution to the existing warning generation methods.
Chenran Li, Aolin Xu 0002, Enna Sachdeva, Teruhisa Misu, Behzad Dariush
ICRA5
2024 Constrained Human-AI Cooperation: An Inclusive Embodied Social Intelligence Challenge
abstract
We introduce Constrained Human-AI Cooperation (CHAIC), an inclusive embodied social intelligence challenge designed to test social perception and cooperation in embodied agents. In CHAIC, the goal is for an embodied agent equipped with egocentric observations to assist a human who may be operating under physical constraints—e.g., unable to reach high places or confined to a wheelchair—in performing common household or outdoor tasks as efficiently as possible. To achieve this, a successful helper must: (1) infer the human's intents and constraints by following the human and observing their behaviors (social perception), and (2) make a cooperative plan tailored to the human partner to solve the task as quickly as possible, working together as a team (cooperative planning). To benchmark this challenge, we create four new agents with real physical constraints and eight long-horizon tasks featuring both indoor and outdoor scenes with various constraints, emergency events, and potential risks. We benchmark planning- and learning-based baselines on the challenge and introduce a new method that leverages large language models and behavior modeling. Empirical evaluations demonstrate the effectiveness of our benchmark in enabling systematic assessment of key aspects of machine social intelligence. Our benchmark and code are publicly available at https://github.com/UMass-Foundation-Model/CHAIC.
Weihua Du, Qiushi Lyu, Jiaming Shan, Zhenting Qi, Sunli Chen, Andi Peng, Tianmin Shu, Kwonjoon Lee, Behzad Dariush, Chuang Gan 0001
NeurIPS10
2024 Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning
abstract
The widespread adoption of commercial autonomous vehicles (AVs) and advanced driver assistance systems (ADAS) may largely depend on their acceptance by society, for which their perceived trustworthiness and interpretability to riders are crucial. In general, this task is challenging because modern autonomous systems software relies heavily on black-box artificial intelligence models. Towards this goal, this paper introduces a novel dataset, Rank2Tell1, a multi-modal ego-centric dataset for Ranking the importance level and Telling the reason for the importance. Using various close and open-ended visual question answering, the dataset provides dense annotations of various semantic, spatial, temporal, and relational attributes of various important objects in complex traffic scenarios. The dense annotations and unique attributes of the dataset make it a valuable resource for researchers working on visual scene understanding and related fields. Furthermore, we introduce a joint model for joint importance level ranking and natural language captions generation to benchmark our dataset and demonstrate performance with quantitative evaluations.
Enna Sachdeva, Nakul Agarwal, Suhas Chundi, Sean Roelofs, Jiachen Li 0001, Mykel J. Kochenderfer, Chiho Choi, Behzad Dariush
WACV8
2023 Weakly-Supervised Action Segmentation and Unseen Error Detection in Anomalous Instructional Videos
abstract
We present a novel method for weakly-supervised action segmentation and unseen error detection in anomalous instructional videos. In the absence of an appropriate dataset for this task, we introduce the Anomalous Toy Assembly (ATA) dataset1, which comprises 1152 untrimmed videos of 32 participants assembling three different toys, recorded from four different viewpoints. The training set comprises 27 participants who assemble toys in an expected and consistent manner, while the test and validation sets comprise 5 participants who display sequential anomalies in their task. We introduce a weakly labeled segmentation algorithm that is a generalization of the constrained Viterbi algorithm and identifies potential anomalous moments based on the difference between future anticipation and current recognition results. The proposed method is not restricted by the training transcripts during testing, allowing for the inference of anomalous action sequences while maintaining real-time performance. Based on these segmentation results, we also introduce a baseline to detect pre-defined human errors, and benchmark results on the ATA dataset. Experiments were conducted on the ATA and CSV datasets, outperforming the state-of-the-art in segmenting anomalous videos under both online and offline conditions.
Reza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Behzad Dariush
ICCV4
2022 Weakly-Supervised Online Action Segmentation in Multi-View Instructional Videos
abstract
This paper addresses a new problem of weakly-supervised online action segmentation in instructional videos. We present a framework to segment streaming videos online at test time using Dynamic Programming and show its advantages over greedy sliding window approach. We improve our framework by introducing the Online-Offline Discrepancy Loss (OODL) to encourage the segmentation results to have a higher temporal consistency. Furthermore, only during training, we exploit framewise correspondence between multiple views as supervision for training weakly-labeled instructional videos. In particular, we investigate three different multi-view inference techniques to generate more accurate frame-wise pseudo ground-truth with no additional annotation cost. We present results and ablation studies on two benchmark multi-view datasets, Breakfast and IKEA ASM. Experimental results show efficacy of the proposed methods both qualitatively and quantitatively in two domains of cooking and assembly.
Reza Ghoddoosian, Isht Dwivedi, Nakul Agarwal, Chiho Choi, Behzad Dariush
CVPR5
2021 Bird's Eye View Segmentation Using Lifted 2D Semantic Features
Isht Dwivedi, Srikanth Malla, Yi-Ting Chen 0001, Behzad Dariush
BMVC4
2021 Social-STAGE: Spatio-Temporal Multi-Modal Future Trajectory Forecast
abstract
This paper considers the problem of multi-modal future trajectory forecast with ranking. Here, multi-modality and ranking refer to the multiple plausible path predictions and the confidence in those predictions, respectively. We propose Social-STAGE, Social interaction-aware Spatio-Temporal multi-Attention Graph convolution network with novel Evaluation for multi-modality. Our main contributions include analysis and formulation of multi-modality with ranking using interaction and multi-attention, and introduction of new metrics to evaluate the diversity and associated confidence of multi-modal predictions. We evaluate our approach on existing public datasets ETH and UCY and show that the proposed algorithm outperforms the state of the arts on these datasets.
Srikanth Malla, Chiho Choi, Behzad Dariush
ICRA3
2020 Unsupervised Domain Adaptation for Spatio-Temporal Action Localization
Nakul Agarwal, Yi-Ting Chen 0001, Behzad Dariush, Ming-Hsuan Yang 0001
BMVC3
2020 TITAN: Future Forecast Using Action Priors
abstract
We consider the problem of predicting the future trajectory of scene agents from egocentric views obtained from a moving platform. This problem is important in a variety of domains, particularly for autonomous systems making reactive or strategic decisions in navigation. In an attempt to address this problem, we introduce TITAN (Trajectory Inference using Targeted Action priors Network), a new model that incorporates prior positions, actions, and context to forecast future trajectory of agents and future ego-motion. In the absence of an appropriate dataset for this task, we created the TITAN dataset that consists of 700 labeled video-clips (with odometry) captured from a moving vehicle on highly interactive urban traffic scenes in Tokyo. Our dataset includes 50 labels including vehicle states and actions, pedestrian age groups, and targeted pedestrian action attributes that are organized hierarchically corresponding to atomic, simple/complex-contextual, transportive, and communicative actions. To evaluate our model, we conducted extensive experiments on the TITAN dataset, revealing significant performance improvement against baselines and state-of-the-art algorithms. We also report promising results from our Agent Importance Mechanism (AIM), a module which provides insight into assessment of perceived risk by calculating the relative influence of each agent on the future ego-trajectory. The dataset is available at https://usa.honda-ri.com/titan.
Srikanth Malla, Behzad Dariush, Chiho Choi
CVPR2
2020 SSP: Single Shot Future Trajectory Prediction
abstract
We propose a robust solution to future trajectory forecast, which can be practically applicable to autonomous agents in highly crowded environments. For this, three aspects are particularly addressed in this paper. First, we use composite fields to predict future locations of all road agents in a singleshot, which results in a constant time complexity, regardless of the number of agents in the scene. Second, interactions between agents are modeled as a non-local response, enabling spatial relationships between different locations to be captured temporally as well (i.e., in spatio-temporal interactions). Third, the semantic context of the scene are modeled and take into account the environmental constraints that potentially influence the future motion. To this end, we validate the robustness of the proposed approach using the ETH, UCY, and SDD datasets and highlight its practical functionality compared to the current state-of-the-art methods.
Isht Dwivedi, Srikanth Malla, Behzad Dariush, Chiho Choi
IROS3
2020 Spatio-Temporal Pyramid Graph Convolutions for Human Action Recognition and Postural Assessment
abstract
Recognition of human actions and associated interactions with objects and the environment is an important problem in computer vision due to its potential applications in a variety of domains. Recently, graph convolutional networks that extract features from the skeleton have demonstrated promising performance. In this paper, we propose a novel Spatio-Temporal Pyramid Graph Convolutional Network (ST-PGN) for online action recognition for ergonomics risk assessment that enables the use of features from all levels of the skeleton feature hierarchy. The proposed algorithm outperforms state-of-art action recognition algorithms tested on two public benchmark datasets typically used for postural assessment (TUM and UW-IOM). We also introduce a pipeline to enhance postural assessment methods with online action recognition techniques. Finally, the proposed algorithm is integrated with a traditional ergonomics risk index (REBA) to demonstrate the potential value for assessment of musculoskeletal disorders in occupational safety.
Behnoosh Parsa, Athma Narayanan, Behzad Dariush
WACV3
2019 Looking to Relations for Future Trajectory Forecast
abstract
Inferring relational behavior between road users as well as road users and their surrounding physical space is an important step toward effective modeling and prediction of navigation strategies adopted by participants in road scenes. To this end, we propose a relation-aware framework for future trajectory forecast. Our system aims to infer relational information from the interactions of road users with each other and with the environment. The first module involves visual encoding of spatio-temporal features, which captures human-human and human-space interactions over time. The following module explicitly constructs pair-wise relations from spatio-temporal interactions and identifies more descriptive relations that highly influence future motion of the target road user by considering its past trajectory. The resulting relational features are used to forecast future locations of the target, in the form of heatmaps with an additional guidance of spatial dependencies and consideration of the uncertainty. Extensive evaluations on the public benchmark datasets demonstrate the robustness and efficacy of the proposed framework as observed by performances higher than the state-of-the-art methods.
Chiho Choi, Behzad Dariush
ICCV2
2019 Dynamic Traffic Scene Classification with Space-Time Coherence
abstract
This paper examines the problem of dynamic traffic scene classification under space-time variations in viewpoint that arise from video captured on-board a moving vehicle. Solutions to this problem are important for realization of effective driving assistance technologies required to interpret or predict road user behavior. Currently, dynamic traffic scene classification has not been adequately addressed due to a lack of benchmark datasets that consider spatiotemporal evolution of traffic scenes resulting from a vehicle's ego-motion. This paper has three main contributions. First, an annotated dataset is released to enable dynamic scene classification that includes 80 hours of diverse high quality driving video data clips collected in the San Francisco Bay area. The dataset includes temporal annotations for road places, road types, weather, and road surface conditions. Second, we introduce novel and baseline algorithms that utilize semantic context and temporal nature of the dataset for dynamic classification of road scenes. Finally, we showcase algorithms and experimental results that highlight how extracted features from scene classification serve as strong priors and help with tactical driver behavior understanding. The results show significant improvement from previously reported driving behavior detection baselines in the literature.
Athma Narayanan, Isht Dwivedi, Behzad Dariush
ICRA3
2019 Egocentric Vision-based Future Vehicle Localization for Intelligent Driving Assistance Systems
abstract
Predicting the future location of vehicles is essential for safety-critical applications such as advanced driver assistance systems (ADAS) and autonomous driving. This paper introduces a novel approach to simultaneously predict both the location and scale of target vehicles in the first-person (egocentric) view of an ego-vehicle. We present a multi-stream recurrent neural network (RNN) encoder-decoder model that separately captures both object location and scale and pixel-level observations for future vehicle localization. We show that incorporating dense optical flow improves prediction results significantly since it captures information about motion as well as appearance change. We also find that explicitly modeling future motion of the ego-vehicle improves the prediction accuracy, which could be especially beneficial in intelligent and automated vehicles that have motion planning capability. To evaluate the performance of our approach, we present a new dataset of first-person videos collected from a variety of scenarios at road intersections, which are particularly challenging moments for prediction because vehicle trajectories are diverse and dynamic. Code and dataset have been made available at: https://usa.honda-ri.com/hevi.
Yu Yao 0006, Chiho Choi, David Crandall, Ella M. Atkins, Behzad Dariush
ICRA6
2019 Ego-motion and Surrounding Vehicle State Estimation Using a Monocular Camera
abstract
Understanding ego-motion and surrounding vehicle state is essential to enable automated driving and advanced driving assistance technologies. Typical approaches to solve this problem use fusion of multiple sensors such as LiDAR, camera, and radar to recognize surrounding vehicle state, including position, velocity, and orientation. Such sensing modalities are overly complex and costly for production of personal use vehicles. In this paper, we propose a novel machine learning method to estimate ego-motion and surrounding vehicle state using a single monocular camera. Our approach is based on a combination of three deep neural networks to estimate the 3D vehicle bounding box, depth, and optical flow from a sequence of images. The main contribution of this paper is a new framework and algorithm that integrates these three networks in order to estimate the ego-motion and surrounding vehicle state. To realize more accurate 3D position estimation, we address ground plane correction in real-time. The efficacy of the proposed method is demonstrated through experimental evaluations that compare our results to ground truth data available from other sensors including Can-Bus and LiDAR.
Jun Hayakawa, Behzad Dariush
IV2
2012 Central mechanisms for force and motion - Towards computational synthesis of human movement
Hooshang Hemami, Behzad Dariush
Neural Networks2
2010 Constrained closed loop inverse kinematics
abstract
This paper introduces a kinematically constrained closed loop inverse kinematics algorithm for motion control of robots or other articulated rigid body systems. The proposed strategy utilizes gradients of collision and joint limit potential functions to arrive at an appropriate weighting matrix to penalize and dampen motion approaching constraint surfaces. The method is particularly suitable for self collision avoidance of highly articulated systems which may have multiple collision points among several segment pairs. In that respect, the proposed method has a distinct advantage over existing gradient projection based methods which rely on numerically unstable null-space projections when there are multiple intermittent constraints. We also show how this approach can be augmented with a previously reported method based on redirection of constraints along virtual surface manifolds. The hybrid strategy is effective, robust, and does not require parameter tuning. The efficacy of the proposed algorithm is demonstrated for a self collision avoidance problem where the reference motion is obtained from human observations. We show simulation and experimental results on the humanoid robot ASIMO.
Behzad Dariush, Youding Zhu, Arjun Arumbakkam, Kikuo Fujimura
ICRA1
2010 Whole-body humanoid control from upper-body task specifications
abstract
This paper introduces a very efficient, modified resolved acceleration control algorithm for dynamic filtering and control of whole-body humanoid motion in response to upper-body task specifications. The dynamic filter is applicable for general upper-body motions when standing in place. It is characterized by modification of the commanded torso acceleration based on a geometric solution to produce a ZMP which is inside the support. The resulting feasible modified motion is synchronized to the reference motion when the computed ZMP for the reference motion again falls within the support. Contact forces at each foot are controlled through a dedicated force distribution module which optimizes the ankle roll and pitch torques. The proposed approach uses time-local information and is therefore targeted for online control. The effectiveness of the algorithm is demonstrated by means of simulated experiments on a model of the Honda humanoid robot ASIMO using a highly dynamic upper-body reference motion.
Ghassan Bin Hammam, David E. Orin, Behzad Dariush
ICRA3
2010 Constrained resolved acceleration control for humanoids
abstract
Resolved acceleration control is a well-known strategy used in tracking control of robotic systems where the desired motion is specified in task-space. Typically, such controllers are developed for systems which exhibit redundancy with respect to execution of operational tasks. While redundancy fundamentally adds new capabilities (self-motion and subtask performance capability), the degree to which secondary objectives can be faithfully executed cannot be determined in advance unless the motion is planned and the environment is known. Therefore, execution of secondary objectives cannot be guaranteed. In fact, a robot which exhibits redundancy with respect to operational tasks may have insufficient degrees of freedom to fulfill more critical objectives such as enforcing constraints. In this paper, we present a generalized constrained resolved acceleration control framework to handle execution of operational tasks and constraints for redundant and non-redundant task (and constraint) specifications. The approach is particularly well suited for online control of complex robot structures such as humanoid robots. The current formulation considers joint limit and collision constraints. The efficacy of the proposed algorithm is demonstrated by simulated experiments of task level upper-body human motion replication on the Honda humanoid robot.
Behzad Dariush, Ghassan Bin Hammam, David E. Orin
IROS1
2010 Kinematic self retargeting: A framework for human pose estimation
Youding Zhu, Behzad Dariush, Kikuo Fujimura
Comput. Vis. Image Underst.2
2009 Toward a vision based hand gesture interface for robotic grasping
abstract
The challenging problem of planning manipulation tasks for dexterous robotic hands can be significantly simplified if the robot system has the ability to learn manipulation skills by observing a human demonstrator. Toward this goal, we present a novel computer vision based hand posture recognition system to serve as an intelligent interface for skill transfer in robotic manipulation. We use the inner distance shape context (IDSC) as a hand shape descriptor to capture variations in the hand state (open or closed) under large in-plane rotations and considerable out-of-plane rotations. The proposed technique is further examined in applications involving grasp recognition and gesture based communications. The experiments show that the proposed approach can be generalized to recognizing a selected taxonomy of grasp types. At present, skin color is used to segment the hand region from the scene, but this method has its own limitations. We show preliminary results suggesting that the IDSC can be used to segment parts of the articulated object, including segmenting the hand from the human body silhouette without using skin color information.
Raghuraman Gopalan, Behzad Dariush
IROS2
2008 Whole body humanoid control from human motion descriptors
abstract
Many advanced motion control strategies developed in robotics use captured human motion data as valuable source of examples to simplify the process of programming or learning complex robot motions. Direct and online control of robots from observed human motion has several inherent challenges. The most important may be the representation of the large number of mechanical degrees of freedom involved in the execution of movement tasks. Attempting to map all such degrees of freedom from a human to a humanoid is a formidable task from an instrumentation and sensing point of view. More importantly, such an approach is incompatible with mechanisms in the central nervous system which are believed to organize or simplify the control of these degrees of freedom during motion execution and motor learning phase. Rather than specifying the desired motion of every degree of freedom for the purpose of motion control, it is important to describe motion by low dimensional motion primitives that are defined in Cartesian (or task) space. In this paper, we formulate the human to humanoid retargeting problem as a task space control problem. The control objective is to track desired task descriptors while satisfying constraints such as joint limits, velocity limits, collision avoidance, and balance. The retargeting algorithm generates the joint space trajectories that are commanded to the robot. We present experimental and simulation results of the retargeting control algorithm on the Honda humanoid robot ASIMO.
Behzad Dariush, Michael Gienger, Bing Jian, Christian Goerick, Kikuo Fujimura
ICRA1
2008 Online and markerless motion retargeting with kinematic constraints
abstract
Transferring motion from a human demonstrator to a humanoid robot is an important step toward developing robots that are easily programmable and that can replicate or learn from observed human motion. The so called motion retargeting problem has been well studied and several off-line solutions exist based on optimization approaches that rely on pre-recorded human motion data collected from a marker-based motion capture system. From the perspective of human robot interaction, there is a growing interest in online and marker-less motion transfer. Such requirements have placed stringent demands on retargeting algorithms and limited the potential use of off-line and pre-recorded methods. To address these limitations, we present an online task space control theoretic retargeting formulation to generate robot joint motions that adhere to the robot’s joint limit constraints, self-collision constraints, and balance constraints. The inputs to the proposed method include low dimensional normalized human motion descriptors, detected and tracked using a vision based feature detection and tracking algorithm. The proposed vision algorithm does not rely on markers placed on anatomical landmarks, nor does it require special instrumentation or calibration. The current implementation requires a depth image sequence, which is collected from a single time of flight imaging device. We present online experimental results of the entire pipeline on the Honda humanoid robot - ASIMO.
Behzad Dariush, Michael Gienger, Arjun Arumbakkam, Christian Goerick, Youding Zhu, Kikuo Fujimura
IROS1
2005 Analysis and Simulation of an Exoskeleton Controller that Accommodates Static and Reactive Loads
abstract
Exoskeletons are structures of rigid links mounted on the body that promise to restore, rehabilitate, or enhance the human motor function. A major challenge in the practical use of exoskeletons for daily activities relate to the coupled control of human-exoskeleton system. This paper provides a method to resolve the control problem by relegation of the human control and exoskeleton control to two control subsystems. The first subsystem represents the execution of voluntary control from commands generated from the central nervous system. This subsystem is responsible primarily for the kinetic or dynamic components of the motion, including motion generation. The second subsystem represents the exoskeleton controller, responsible for joint level accommodation of all gravitational, static, and certain reactive forces. If all such forces are static, the exoskeleton controller can be viewed as a compensator that maintains the body in static equilibrium. The proposed strategy provides a clear partition between natural voluntary control by the CNS, and artificial assist by the exoskeleton controller. Two methods are presented for implementation of the proposed control algorithm. The first method is based on the principal of virtual work. The second method is a recursive algorithm based on force and moment balance equations. Analytical results are presented to study feasibility regions of exoskeleton control strategies in terms of mechanical efficiency. Finally, simulation results are presented to demonstrate the efficacy of the algorithm for a powered ankle-foot orthosis application.
Behzad Dariush
ICRA1
2003 Human motion analysis for biomechanics and biomedicine
Behzad Dariush
Mach. Vis. Appl.1
2000 Analysis and Synthesis of Human Motion from External Measurements
abstract
The structures observed in humans are being progressively applied to the theoretical approaches developed in robotics. To gain insight to the intricate mechanism of human motion, researchers sometimes use imaging technology to record the trajectories of humans performing various tasks. From these observations, they are able to estimate the forces and moments at each joint by an inverse dynamics computation. This problem is conceptually simple; however, in practice, the inverse solution requires the calculation of higher order derivatives of experimental observations contaminated by noise. The errors due to differentiation results in erroneous joint force and moment calculations. This paper provide a control theoretic framework for analyzing human motion which avoids derivative computations. The method is also suitable for synthesis of stable controllers for robotic and 'biorobic' applications which require tracking a desired reference trajectory under different loading conditions.
Behzad Dariush, Hooshang Hemami, Mohamad Parnianpour
ICRA1
2000 Single Rigid Body Representation, Control and Stability for Robotic Applications
abstract
In this paper, a novel formulation of the dynamics, control and stability of a single rigid body is presented. Effects of gravity, and visco-elastic coupling to an inertial frame of reference at a single point of contact are included. This formulation is very convenient, and analytically tractable for animation and computer simulation of human, animal, robotic and humanoid movements, for studies of performance assessment and enhancement in natural and man-made systems, and in other studies of systems of connected rigid bodies. The representation includes simple and general linear and nonlinear position and velocity, i.e., state feedback structures that guarantee asymptotic stability of the system in the sense of Lyapunov. The Lyapunov function is simply the physical energy stored in the system, namely, a quadratic in the state space of the system: the sum of kinetic, elastic and potential energies of the system The approach is extended to a rigid body coupled to an inertial frame of reference. Digital computer simulation of the behavior of the system under disturbance are presented.
Hooshang Hemami, Behzad Dariush
ICRA2
1998 Spatiotemporal Analysis of Face Profiles: Detection, Segmentation, and Registration
Behzad Dariush, Sing Bang Kang, Keith Waters
FG1