EDBT 2026 Demo / reviewers in the wild / expert
Tariq Iqbal
dblp:159/0463
· DBLP profile ↗
30ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 5 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 14 · 3 first-author · 11 since 2021Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robot-Assisted Medical Training for Safety-Critical EnvironmentsabstractWhile resuscitation training is critical, healthcare workers (HCWs) with high workload have limited chance to get trained and re-trained due to time and resource constraints. To address this gap, we engaged in a co-design process of robots that facilitate and prepare HCWs for resuscitation procedures (i.e., codes). First, we investigated what resuscitation training consists of, including challenges faced by trainees and trainers. Second, we collaboratively explored how a crash cart robot, that guides users to medical supplies and equipment, could assist trainers and trainees synchronously–during team-based clinical simulations and asynchronously–during one-on-one training. We found that robots could 1) serve as a learning assistant by providing real-time feedback and supporting personalized training needs; and 2) an evaluating assistant by monitoring multiple trainees and tracking critical timing of interventions in the training. Through this new training paradigm, we hope to demonstrate opportunities for crash cart robots to aid HCWs for their sustainable training and reskilling. We discuss the role of robots in training beyond cognitive knowledge, situating them within two underexplored contexts: practical skill training and team-based training. Huajie Cao, Michael J. Sack, Lili Mkrtchyan, Kevin Ching, Tariq Iqbal, Hee Rin Lee, Angelique Taylor |
HRI | 5 |
| 2026 | Embodied Referring Expression Comprehension in Human-Robot InteractionabstractAs robots enter human workspaces, there is a crucial need for them to comprehend embodied human instructions, enabling intuitive and fluent human-robot interaction (HRI). However, accurate comprehension is challenging due to a lack of large-scale datasets that capture natural embodied interactions in diverse HRI settings. Existing datasets suffer from perspective bias, single-view data collection, inadequate coverage of nonverbal gestures, and a predominant focus on indoor environments. To address these issues, we present the Refer360 dataset, a large-scale dataset of embodied verbal and nonverbal interactions collected across diverse viewpoints in both indoor and outdoor settings. Additionally, we introduce MuRes, a multimodal guided residual module designed to improve embodied referring expression comprehension. MuRes acts as an information bottleneck, extracting salient modality-specific signals and reinforcing them into pre-trained representations to form complementary features for downstream tasks. We conduct extensive experiments on four datasets, including our Refer360 dataset, and demonstrate that current multimodal models fail to capture embodied interactions comprehensively; however, augmenting them with MuRes consistently improves performance. These findings establish Refer360 as a valuable benchmark and exhibit the potential of guided residual learning to advance embodied referring expression comprehension in robots operating within human environments. Md. Mofijul Islam, Alexi Gladstone, Sujan Sarker, Ganesh Nanduru, Md Fahim, Keyan Du, Aman Chadha, Tariq Iqbal |
HRI | 8 |
| 2026 | Examining Physiological Response and Facial Expression as Indicators of Trust in a Robot PartnerabstractDevelopments in the assistive capabilities of robots have allowed them to be more prevalent across domains from manufacturing to healthcare to education. To achieve more efficient and fluent collaborations between robots and humans, it has become necessary to explore how humans trust robot partners during collaborations. Current measurements of human trust in a robot partner typically rely on reflective, subjective survey measures, which are difficult to incorporate and update in the robot’s decision-making models for real-time deployment. To address this challenge, we aim to identify potential objective measures of trust, which can be incorporated into the robot’s decision-making process. We examine physiological data (Electrodermal Activity (EDA), Skin Temperature (TEMP), and Blood Volume Pulse (BVP)) and facial expression in an in-person ( \(n=30\) ) human–robot collaboration study and report the effect of trust on these objective measures. The results indicate that the participants who distrusted the robot partner experienced higher EDA, higher TEMP, and lower BVP variations than participants who trusted the robot. For facial expression, participants who distrusted the robot had more intense facial expressions than participants who trusted the robot. Ultimately, these findings can be used to further research on developing a real-time model of human trust in robots. Haley N. Green, Tariq Iqbal |
ACM Trans. Hum. Robot Interact. | 2 |
| 2026 | A Holistic Evaluation of Teleoperation Interfaces for Robotic ManipulationabstractRobot teleoperation has become increasingly crucial for extending human capabilities in inaccessible or hazardous environments and facilitating human–robot collaboration. While significant advancements have been made in teleoperation interfaces, the success of these systems critically depends on how effectively humans can interact with and control robotic systems across diverse manipulation tasks. However, existing research primarily evaluates interfaces within specific tasks or applications, lacking systematic assessment across different manipulation scenarios. This limitation leads to suboptimal interface selection that can compromise task efficiency in critical domains and impede the collection of high-quality demonstrations for robot learning. To address these gaps, we first introduce a novel two-axis Robotic Manipulation Task Taxonomy that systematically categorizes manipulation tasks based on their fundamental control requirements: Motion Type (translation-dominant vs. rotation-dominant) and Engagement Type (rigid vs. non-rigid object interactions). We conducted a comprehensive user study ( \(n=30\) ) evaluating three distinct teleoperation interfaces (Gamepad, 3D Mouse, and Virtual Reality (VR) Controller) based on this taxonomy. Our results indicate that the Gamepad and 3D mouse significantly outperformed the VR Controller in task completion time across all task categories. In contrast, the VR Controller showed higher first-attempt success rates but caused significantly greater cognitive load across all NASA Task Load Index (NASA-TLX) subscales compared to the Gamepad interface, and specifically higher mental demand, physical demand, and effort compared to the 3D mouse. Our results also indicate that although prior familiarity with an interface lowered perceived workload and enhanced perceived performance, it did not translate into improvements in actual task success rates or completion times. These insights provide valuable guidelines for optimizing teleoperation interfaces based on task requirements and highlight the importance of considering both cognitive demands and user experience in interface design. Shaid Hasan, Mohammad Samin Yasar, Tariq Iqbal |
ACM Trans. Hum. Robot Interact. | 3 |
| 2025 | Using Physiological Measures, Gaze, and Facial Expressions to Model Human Trust in a Robot PartnerabstractWith robots becoming increasingly prevalent in various domains, it has become crucial to equip them with tools to achieve greater fluency in interactions with humans. One of the promising areas for further exploration lies in human trust. A real-time, objective model of human trust could be used to maximize productivity, preserve safety, and mitigate failure. In this work, we attempt to use physiological measures, gaze, and facial expressions to model human trust in a robot partner. We are the first to design an inperson, human-robot supervisory interaction study to create a dedicated trust dataset. Using this dataset, we train machine learning algorithms to identify the objective measures that are most indicative of trust in a robot partner, advancing trust prediction in human-robot interactions. Our findings indicate that a combination of sensor modalities (blood volume pulse, electrodermal activity, skin temperature, and gaze) can enhance the accuracy of detecting human trust in a robot partner. Furthermore, the Extra Trees, Random Forest, and Decision Trees classifiers exhibit consistently better performance in measuring the person's trust in the robot partner. These results lay the groundwork for constructing a real-time trust model for human-robot interaction, which could foster more efficient interactions between humans and robots. Haley N. Green, Tariq Iqbal |
ICRA | 2 |
| 2024 | PoseTron: Enabling Close-Proximity Human-Robot Collaboration Through Multi-human Motion PredictionabstractAs robots enter human workspaces, there is a crucial need for robots to understand and predict human motion to achieve safe and fluent human-robot collaboration (HRC). However, accurate prediction is challenging due to a lack of large-scale datasets for close-proximity HRC and the absence of generalizable algorithms. To overcome these challenges, we present INTERACT, a comprehensive multimodal dataset covering 3-D Skeleton, RGB+D, gaze, and robot joint data for human-human and human-robot collaboration. Additionally, we introduce PoseTron, a novel transformer-based architecture to address the gap in learning algorithms. PoseTron introduces a conditional attention mechanism in the encoder enabling efficient weighing of motion information from all agents to incorporate team dynamics. The decoder features a novel multimodal attention mechanism, which weights representations from different modalities and the encoder outputs to predict future motion. We extensively evaluated PoseTron by comparing its performance on the INTERACT dataset against state-of-the-art algorithms. The results suggest that PoseTron outperformed all other methods across all the scenarios, attaining lowest prediction errors. Furthermore, we conducted a comprehensive ablation study, emphasizing the importance of design choices, pointing towards a promising direction for integrating motion prediction with robot perception in safe and effective HRC. Mohammad Samin Yasar, Md. Mofijul Islam, Tariq Iqbal |
HRI | 3 |
| 2024 | EQA-MX: Embodied Question Answering using Multimodal ExpressionabstractHumans predominantly use verbal utterances and nonverbal gestures (e.g., eye gaze and pointing gestures) in their natural interactions. For instance, pointing gestures and verbal information is often required to comprehend questions such as "what object is that?" Thus, this question-answering (QA) task involves complex reasoning of multimodal expressions (verbal utterances and nonverbal gestures). However, prior works have explored QA tasks in non-embodied settings, where questions solely contain verbal utterances from a single verbal and visual perspective. In this paper, we have introduced 8 novel embodied question answering (EQA) tasks to develop learning models to comprehend embodied questions with multimodal expressions. We have developed a novel large-scale dataset, EQA-MX, with over 8 million diverse embodied QA data samples involving multimodal expressions from multiple visual and verbal perspectives. To learn salient multimodal representations from discrete verbal embeddings and continuous wrapping of multiview visual representations, we propose a vector-quantization (VQ) based multimodal representation learning model, VQ-Fusion, for the EQA tasks. Our extensive experimental results suggest that VQ-Fusion can improve the performance of existing state-of-the-art visual-language models up to 13% across EQA tasks. Md. Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq Iqbal |
ICLR | 4 |
| 2024 | M2RL: A Multimodal Multi-Interface Dataset for Robot Learning from Human DemonstrationsabstractImitation Learning, inspired by observational learning theory in cognitive psychology, is a promising approach for teaching robots to perform complex manipulation tasks. However, most imitation learning datasets exhibit biases by focusing on a single interface or modality when capturing human demonstrations. This limitation fails to fully capture the multimodal nature of how humans learn skills through demonstration. To bridge this gap, we introduce the M2RL dataset, a multimodal and multi-interface dataset collected from non-expert users across diverse manipulation tasks from four task categories using three distinct teleoperation interfaces. The M2RL dataset comprises RGB+D data from three camera perspectives (robot’s wrist and two exo-views), ego-view and gaze data from the human teleoperator’s perspective, and the robot’s proprioception data. Our extensive evaluation of state-of-the-art imitation learning algorithms on the M2RL dataset highlights the importance of multimodal and multi-interface data for learning robust policies for the robot. Additionally, the results indicate clear performance improvements when training on data from diverse interfaces and utilizing inputs from multiple camera streams. Our dataset and code are publicly available at: https://github.com/M2RL/m2rl-dataset. Shaid Hasan, Mohammad Samin Yasar, Tariq Iqbal |
ICMI | 3 |
| 2024 | What Am I? Evaluating the Effect of Language Fluency and Task Competency on the Perception of a Social RobotabstractRecent advancements in robot capabilities have enabled them to interact with people in various human-social environments (HSEs). In many of these environments, the perception of the robot often depends on its capabilities, e.g., task competency, language fluency, etc. To enable fluent human-robot interaction (HRI) in HSEs, it is crucial to understand the impact of these capabilities on the perception of the robot. Although many works have investigated the effects of various robot capabilities on the robot’s perception separately, in this paper, we present a large-scale HRI study (n = 60) to investigate the combined impact of both language fluency and task competency on the perception of a robot. The results suggest that while language fluency may play a more significant role than task competency in the perception of the verbal competency of a robot, both language fluency and task competency contribute to the perception of the intelligence and reliability of the robot. The results also indicate that task competency may play a more significant role than language fluency in the perception of meeting expectations and being a good teammate. The findings of this study highlight the relationship between language fluency and task competency in the context of social HRI and will enable the development of more intelligent robots in the future. Shahira Ali, Haley N. Green, Tariq Iqbal |
RO-MAN | 3 |
| 2024 | IMPRINT: Interactional Dynamics-aware Motion Prediction in Teams using Multimodal ContextabstractRobots are moving from working in isolation to working with humans as a part of human-robot teams. In such situations, they are expected to work with multiple humans and need to understand and predict the team members’ actions. To address this challenge, in this work, we introduce IMPRINT, a multi-agent motion prediction framework that models the interactional dynamics and incorporates the multimodal context (e.g., data from RGB and depth sensors and skeleton joint positions) to accurately predict the motion of all the agents in a team. In IMPRINT, we propose an Interaction module that can extract the intra-agent and inter-agent dynamics before fusing them to obtain the interactional dynamics. Furthermore, we propose a Multimodal Context module that incorporates multimodal context information to improve multi-agent motion prediction. We evaluated IMPRINT by comparing its performance on human-human and human-robot team scenarios against state-of-the-art methods. The results suggest that IMPRINT outperformed all other methods over all evaluated temporal horizons. Additionally, we provide an interpretation of how IMPRINT incorporates the multimodal context information from all the modalities during multi-agent motion prediction. The superior performance of IMPRINT provides a promising direction to integrate motion prediction with robot perception and enable safe and effective human-robot collaboration. Mohammad Samin Yasar, Md. Mofijul Islam, Tariq Iqbal |
ACM Trans. Hum. Robot Interact. | 3 |
| 2023 | PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal CuesabstractHumans naturally use referring expressions with verbal utterances and nonverbal gestures to refer to objects and events. As these referring expressions can be interpreted differently from the speaker's or the observer's perspective, people effectively decide on the perspective in comprehending the expressions. However, existing models do not explicitly learn perspective grounding, which often causes the models to perform poorly in understanding embodied referring expressions. To make it exacerbate, these models are often trained on datasets collected in non-embodied settings without nonverbal gestures and curated from an exocentric perspective. To address these issues, in this paper, we present a perspective-aware multitask learning model, called PATRON, for relation and object grounding tasks in embodied settings by utilizing verbal utterances and nonverbal cues. In PATRON, we have developed a guided fusion approach, where a perspective grounding task guides the relation and object grounding task. Through this approach, PATRON learns disentangled task-specific and task-guidance representations, where task-guidance representations guide the extraction of salient multimodal features to ground the relation and object accurately. Furthermore, we have curated a synthetic dataset of embodied referring expressions with multimodal cues, called CAESAR-PRO. The experimental results suggest that PATRON outperforms the evaluated state-of-the-art visual-language models. Additionally, the results indicate that learning to ground perspective helps machine learning models to improve the performance of the relation and object grounding task. Furthermore, the insights from the extensive experimental results and the proposed dataset will enable researchers to evaluate visual-language models' effectiveness in understanding referring expressions in other embodied settings. Md. Mofijul Islam, Alexi Gladstone, Tariq Iqbal |
AAAI | 3 |
| 2023 | Representation Learning in Deep RL via Discrete Information BottleneckabstractSeveral self-supervised representation learning methods have been proposed for reinforcement learning (RL) with rich observations. For real world applications of RL, recovering underlying latent states is crucial, particularly when sensory inputs can contain irrelevant and exogenous information. In this work, we study how information bottlenecjs can be used to construct latent states efficiently in the presence of task irrelevant information. We propose architectures that utilize variational and discrete information bottleneck, coined as RepDIB, to learn structured factorized representations. Exploiting the expressiveness bought by factorized representations, we introduce a simple, yet effective, bottleneck that can be integrated with any existing self supervised objective for RL. We demonstrate this across several online and offline RL benchmarks, along with a real robot arm task, where we find that compressed representations with RepDIB can lead to strong performance improvements, as the learnt bottlenecks can help predict only the relevant state, while ignoring irrelevant information. Riashat Islam, Hongyu Zang, Manan Tomar, Aniket Didolkar, Md. Mofijul Islam, Samin Yeasar Arnob, Tariq Iqbal, Xin Li 0033, Anirudh Goyal, Nicolas Heess, Alex Lamb |
AISTATS | 7 |
| 2023 | Reimagining Robots for Dementia: From Robots for Care-receivers/giver to Robots for CarepartnersabstractInformal caregivers are the main source of dementia care. Considering the importance of both family caregivers and persons living with dementia (PwDs), this paper explores how these two parties go through their dementia journey and how they envision robots to support them. We adopt a person-centered care approach which views these couples as reciprocal carepartners, rather than as care-givers and care-receivers. We conducted a community-based participatory research study with a dementia advocacy organization to imagine how robots can support these dementia dyads. The contribution of this paper is threefold: First, we introduce a person-centered care approach and show how this new approach reveals the issues of PwDs and carepartners (CPs) as partners and citizens. For example, PwDs' main challenges were not dementia symptoms but the concomitant stigma such as fears of being considered abnormal. This issue has rarely been discussed in HRI. Second, we suggest slow communication as an important robot design feature. When robots can wait for PwDs to proceed with information without judging PwDs' relatively slow response, PwDs feel respected and less stigmatized. Third, we address the importance of paying attention to disagreements between PwDs and CPs about robot design preferences. Considering the interdependency of the two parties, robot design processes should allow the two to negotiate. Hee Rin Lee, Tariq Iqbal, Brenda Roberts |
HRI | 3 |
| 2023 | VADER: Vector-Quantized Generative Adversarial Network for Motion PredictionabstractHuman motion prediction is an essential component for enabling close-proximity human-robot collaboration. The task of accurately predicting human motion is non-trivial and is compounded by the variability of human motion and the presence of multiple humans in proximity. To address some of the open challenges in motion prediction, in this work, we propose VADER, a novel sequence learning algorithm that models past observed poses using a flexible discrete latent space. VADER introduces the concept of Vector Quantization for human motion prediction, enabling the learning of a discrete latent space without being restricted by any static prior. In addition, we propose a new objective function that uses the discriminator objective to penalize deviation of predicted motion from the ground-truth. Finally, to explicitly model interaction in multiple humans, we introduce a lightweight attention mechanism to condition per-agent prediction on the previous hidden states of all the agents. Our evaluation across three scenarios: single-agent, multi-agent, and human-robot collaboration shows that VADER outperformed all the state-of-the-art approaches, resulting in more feasible human poses that align better with the ground-truth. Finally, we conducted extensive ablation studies to emphasize the importance of the proposed modules. Mohammad Samin Yasar, Tariq Iqbal |
IROS | 2 |
| 2023 | Properties and estimation approaches of the odd JCA family with applicationsabstractSummary In this article, we study a new generator of distributions called the odd JCA‐G family. We determine the main mathematical properties of the new family. Some special submodels of the odd JCA‐G family are presented. The special models of the JCA‐G family have tractable densities shapes which possess various kinds of asymmetric, reversed‐J, left‐skewed, right‐skewed, symmetrical shapes. Furthermore, special submodels exhibit flexible hazard rate shapes. Additionally, the parameters of the odd JCA Burr‐XII model are estimated using some classical estimation approaches. The performance and efficiency of these approaches are explored via numerical simulations. The applicability of the odd JCA‐G family is illustrated by fitting two real‐life datasets. The data analysis shows that the odd JCA Burr‐XII model outperforms some well‐known distributions for the selected data. Tariq Iqbal, Nada M. Alfaer, Muhammad Hussain Tahir, Hassan M. Aljohani, Farrukh Jamal, Ahmed Z. Afify |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | MAVEN: A Memory Augmented Recurrent Approach for Multimodal FusionabstractMultisensory systems provide complementary information that aids many machine learning approaches in perceiving the environment comprehensively. These systems consist of heterogeneous modalities, which have disparate characteristics and feature distributions. Thus, extracting, aligning, and fusing complementary representations from heterogeneous modalities (e.g., visual, skeleton, and physical sensors) remains challenging. To address these challenges, we have used the insights from several neuroscience studies of animal multisensory systems to develop MAVEN, a memory-augmented recurrent approach for multimodal fusion. MAVEN generates unimodal memory banks comprised of spatial-temporal features and uses our proposed recurrent representation alignment approach to align and refine unimodal representations iteratively. MAVEN then utilizes a multimodal variational attention-based fusion approach to produce a robust multimodal representation from the aligned unimodal features. Our extensive experimental evaluations on three multimodal datasets suggest that MAVEN outperforms state-of-the-art multimodal learning approaches in the challenging human activity recognition task across all evaluation conditions (cross-subject, leave-one-subject-out, and cross-session). Additionally, our extensive ablation studies suggest that MAVEN significantly outperforms the feed-forward fusion-based learning models$(p< 0.05)$. Finally, the robust performance of MAVEN in extracting complementary multimodal representation from occluded and noisy data suggests its applicability on real-world datasets. Md. Mofijul Islam, Mohammad Samin Yasar, Tariq Iqbal |
IEEE Trans. Multim. | 3 |
| 2022 | MuMu: Cooperative Multitask Learning-Based Guided Multimodal FusionabstractMultimodal sensors (visual, non-visual, and wearable) can provide complementary information to develop robust perception systems for recognizing activities accurately. However, it is challenging to extract robust multimodal representations due to the heterogeneous characteristics of data from multimodal sensors and disparate human activities, especially in the presence of noisy and misaligned sensor data. In this work, we propose a cooperative multitask learning-based guided multimodal fusion approach, MuMu, to extract robust multimodal representations for human activity recognition (HAR). MuMu employs an auxiliary task learning approach to extract features specific to each set of activities with shared characteristics (activity-group). MuMu then utilizes activity-group-specific features to direct our proposed Guided Multimodal Fusion Approach (GM-Fusion) for extracting complementary multimodal representations, designed as the target task. We evaluated MuMu by comparing its performance to state-of-the-art multimodal HAR approaches on three activity datasets. Our extensive experimental results suggest that MuMu outperforms all the evaluated approaches across all three datasets. Additionally, the ablation study suggests that MuMu significantly outperforms the baseline models (p Md. Mofijul Islam, Tariq Iqbal |
AAAI | 2 |
| 2022 | Who's Laughing NAO?: Examining Perceptions of Failure in a Humorous Robot PartnerabstractSocial robots are being deployed to interact with people in various scenarios, where they are expected to in-corporate human-like conversational strategies to achieve flu-ency in interactions. For example, current robots are designed to perform advanced communication strategies (i.e., personal anecdotes, explanations, and apologies) to recover from task failure. However, these tactics are not always sufficient for failure recovery as they can be lengthy and insufficient for encouraging future interactions. In human-human interactions, people often use humor as a low-risk and engaging method for managing failures. Thus, the successful execution of advanced, human-like humor could enable robots to recover from task failures more efficiently. In this paper, we present a human-robot interaction study exploring how a robot's utilization of various human-like humor types (i.e., affiliative, aggressive, self-enhancing, and self-defeating) are perceived by human teammate (n = 32) and an external observer of the interaction (n = 256). Additionally, we have explored the effects of performance, humor type, perspective, and previous experience with robots on the participants' perceptions of warmth, competence, and the robot as a teammate. Our results indicate that dyadic participants rated the successful robot to be more competent and a better teammate than the bystander participants. Additionally, the results indicate that participants with less experience with robots found the successful robot to be more competent than participants with high levels of experience. These findings will enable the human-robot interaction community to develop more engaging robots for fluent interactive experiences in the future. Haley N. Green, Md. Mofijul Islam, Shahira Ali, Tariq Iqbal |
HRI | 4 |
| 2022 | Robots That Can Anticipate and Learn in Human-Robot TeamsabstractRobots are moving from working in isolated cham-bers to working in close-proximity with human collaborator(s) as part of human-robot teams. In such situations, robots are increasingly expected to work with multiple humans and ef-fectively model both human-human and human-robot dynamics before taking timely actions. Working toward this goal, we have proposed new algorithms that model human intent and motion while being interpretable and scalable to multiple humans. Our current work builds upon these algorithms to 1) obtain a more holistic representation of the environment and 2) interleave robot perception and control. Our proposed algorithms have attained state-of-the-art performances over various benchmarks and learning scenarios. As part of future work, we aim to enhance our learning algorithms with the capability of acquiring knowledge continually, without overwriting past information. Mohammad Samin Yasar, Tariq Iqbal |
HRI | 2 |
| 2022 | CAESAR: An Embodied Simulator for Generating Multimodal Referring Expression DatasetsabstractHumans naturally use verbal utterances and nonverbal gestures to refer to various objects (known as $\textit{referring expressions}$) in different interactional scenarios. As collecting real human interaction datasets are costly and laborious, synthetic datasets are often used to train models to unambiguously detect relationships among objects. However, existing synthetic data generation tools that provide referring expressions generally neglect nonverbal gestures. Additionally, while a few small-scale datasets contain multimodal cues (verbal and nonverbal), these datasets only capture the nonverbal gestures from an exo-centric perspective (observer). As models can use complementary information from multimodal cues to recognize referring expressions, generating multimodal data from multiple views can help to develop robust models. To address these critical issues, in this paper, we present a novel embodied simulator, CAESAR, to generate multimodal referring expressions containing both verbal utterances and nonverbal cues captured from multiple views. Using our simulator, we have generated two large-scale embodied referring expression datasets, which we have released publicly. We have conducted experimental analyses on embodied spatial relation grounding using various state-of-the-art baseline models. Our experimental results suggest that visual perspective affects the models' performance; and that nonverbal cues improve spatial relation grounding accuracy. Finally, we will release the simulator publicly to allow researchers to generate new embodied interaction datasets. Md. Mofijul Islam, Reza Mirzaiee, Alexi Gladstone, Haley N. Green, Tariq Iqbal |
NeurIPS | 5 |
| 2022 | Fusing Computer Vision and Wireless Signal for Accurate Sensor Localization in AR ViewabstractRecent years have seen increasing traction to enable new applications that can localize sensors on the screen of an Augmented Reality (AR) device (e.g. smartphone, tablet) so that sensors can be controlled more intuitively. Despite recent advances in this area, both wireless signal dependent and computer vision based localization solutions have seen a slow acceptance due to signal noise, multipath effect, and limited AR device-sensor interactivity. In this paper, we propose a novel solution to combine the complementary advantages of wireless signal based localization solution with the computer vision based solution to track IoT devices and sensors more accurately. Experimental result shows that our system can accurately track IoT devices with an average pixel error of 34 pixels in a 1024 × 768 pixels image, which is a 75.8% improvement from the state-of-the-art model. Md Fazlay Rabbi Masum Billah, Md. Mofijul Islam, Nurani Saoda, Tariq Iqbal, Bradford Campbell |
SenSys | 4 |
| 2021 | Temporal Anticipation and Adaptation Methods for Fluent Human-Robot TeamingabstractAs robots work with human teams, they will be expected to fluently coordinate with them. While people are adept at coordination and real-time adaptation, robots still lack this skill. In this paper, we introduce TANDEM: Temporal Anticipation and Adaptation for Machines, a series of neurobiologically-inspired algorithms that enable robots to fluently coordinate with people. TANDEM leverages a humanlike understanding of external and internal temporal changes to facilitate coordination. We experimentally validated the approach via a human-robot collaborative drumming task across tempo-changing rhythmic conditions. We found that an adaptation process alone enables a robot to achieve human-level performance. Moreover, by combining anticipatory knowledge along with an adaptation process, robots can potentially perform such tasks better than people. We hope this work will enable researchers to create robots more sensitive to changes in team dynamics. Tariq Iqbal, Laurel D. Riek |
ICRA | 1 |
| 2020 | HAMLET: A Hierarchical Multimodal Attention-based Human Activity Recognition AlgorithmabstractTo fluently collaborate with people, robots need the ability to recognize human activities accurately. Although modern robots are equipped with various sensors, robust human activity recognition (HAR) still remains a challenging task for robots due to difficulties related to multimodal data fusion. To address these challenges, in this work, we introduce a deep neural network-based multimodal HAR algorithm, HAMLET. HAMLET incorporates a hierarchical architecture, where the lower layer encodes spatio-temporal features from unimodal data by adopting a multi-head self-attention mechanism. We develop a novel multimodal attention mechanism for disentangling and fusing the salient unimodal features to compute the multimodal features in the upper layer. Finally, multimodal features are used in a fully connect neural-network to recognize human activities. We evaluated our algorithm by comparing its performance to several state-of-the-art activity recognition algorithms on three human activity datasets. The results suggest that HAMLET outperformed all other evaluated baselines across all datasets and metrics tested, with the highest top-1 accuracy of 95.12% and 97.45% on the UTD-MHAD [1] and the UT-Kinect [2] datasets respectively, and F1-score of 81.52% on the UCSD-MIT [3] dataset. We further visualize the unimodal and multimodal attention maps, which provide us with a tool to interpret the impact of attention mechanisms concerning HAR. Md. Mofijul Islam, Tariq Iqbal |
IROS | 2 |
| 2019 | Fast Online Segmentation of Activities from Partial TrajectoriesabstractAugmenting a robot with the capacity to understand the activities of the people it collaborates with in order to then label and segment those activities allows the robot to generate an efficient and safe plan for performing its own actions. In this work, we introduce an online activity segmentation algorithm that can detect activity segments by processing a partial trajectory. We model the transitions through activities as a hidden Markov model, which runs online by implementing an efficient particle-filtering approach to infer the maximum a posteriori estimate of the activity sequence. This process is complemented by an online search process to refine activity segments using task model information about the partial order of activities. We evaluated our algorithm by comparing its performance to two state-of-the-art activity segmentation algorithms on three human activity datasets. The proposed algorithm improved activity segmentation accuracy across all three datasets compared with the other two approaches, with a range from 11.3% to 65.5%, and could accurately recognize an activity through observation alone for 31.6% of the initial trajectory of that activity, on average. We also implemented the algorithm onto an industrial mobile robot during an automotive assembly task in which the robot tracked a human worker's progress and provided the worker with the correct materials at the appropriate time. Tariq Iqbal, Shen Li 0003, Christopher K. Fourie, Bradley Hayes, Julie A. Shah |
ICRA | 1 |
| 2019 | Activity recognition in manufacturing: The roles of motion capture and sEMG+inertial wearables in detecting fine vs. gross motionabstractIn safety-critical environments, robots need to reliably recognize human activity to be effective and trust-worthy partners. Since most human activity recognition (HAR) approaches rely on unimodal sensor data (e.g. motion capture or wearable sensors), it is unclear how the relationship between the sensor modality and motion granularity (e.g. gross or fine) of the activities impacts classification accuracy. To our knowledge, we are the first to investigate the efficacy of using motion capture as compared to wearable sensor data for recognizing human motion in manufacturing settings. We introduce the UCSD-MIT Human Motion dataset, composed of two assembly tasks that entail either gross or fine-grained motion. For both tasks, we compared the accuracy of a Vicon motion capture system to a Myo armband using three widely used HAR algorithms. We found that motion capture yielded higher accuracy than the wearable sensor for gross motion recognition (up to 36.95%), while the wearable sensor yielded higher accuracy for fine-grained motion (up to 28.06%). These results suggest that these sensor modalities are complementary, and that robots may benefit from systems that utilize multiple modalities to simultaneously, but independently, detect gross and fine-grained motion. Our findings will help guide researchers in numerous fields of robotics including learning from demonstration and grasping to effectively choose sensor modalities that are most suitable for their applications. Alyssa Kubota, Tariq Iqbal, Julie A. Shah, Laurel D. Riek |
ICRA | 2 |
| 2016 | Human Coordination Dynamics with Heterogeneous Robots in a TeamabstractRobots with different behaviors will be a part of human-robot teams in the future and will impact the overall interaction patterns of teams. In this paper, we investigate how the presence of robots affect the coordination of human-robot teams when a single robot or multiple robots with the same or different behavior are the part of that team. We compare two different event anticipation methods for robots, and then extend those findings to assess its effects on the group coordination. Our results indicate that humans are significantly more synchronous as a group when they danced alone than with the robots. We also find that an addition of a robot with a different anticipation algorithm to a single robot team significantly reduces the group synchrony. This work will prove useful for the robotics community to build more fluent human-robot interactions in the future. Tariq Iqbal, Laurel D. Riek |
HRI | 1 |
| 2016 | A Method for Automatic Detection of Psychomotor EntrainmentabstractGroup interaction is an important aspect of human social behavior. During some group events, the activities performed by each group member continually influence the activities of others. This process of influence can lead to synchronized group activity, or the entrainment of the group. Understanding entrainment is important, because it can be a critical behavioral indicator of group cohesiveness, and can provide context for accurately understanding a group's affective behavior. In this paper, we present a novel method to automatically detect group psychomotor entrainment, which takes multiple types of discrete, task-level events into consideration. We experimentally validated the method on two synchronous rhythmic activities, “the cup game” and a marching task. We also compared its accuracy against two alternate synchrony detection methods in the literature. The results suggest our method can successfully measure group psychomotor entrainment, and is more accurate compared to other methods. This method will be useful to researchers interested in quantitatively and automatically measuring entrainment, and can also provide insight into understanding how groups interact socially. Tariq Iqbal, Laurel D. Riek |
IEEE Trans. Affect. Comput. | 1 |
| 2016 | Movement Coordination in Human-Robot Teams: A Dynamical Systems ApproachabstractIn order to be effective teammates, robots need to be able to understand high-level human behavior to recognize, anticipate, and adapt to human motion. We have designed a new approach to enable robots to perceive human group motion in real time to anticipate future actions and synthesize their own motion accordingly. We explore this within the context of joint action, in which humans and robots move together synchronously. In this paper we present an anticipation method, which takes high-level group behavior into account. We validate the method within a human-robot interaction scenario, in which an autonomous mobile robot observes a team of human dancers and then successfully and contingently coordinates its movements to “join the dance.” We compared the results of our anticipation method to move the robot with another method that did not rely on high-level group behavior and found that our method performed better both in terms of more closely synchronizing the robot's motion to the team and exhibiting more contingent and fluent motion. These findings suggest that the robot performs better when it has an understanding of high-level group behavior than when it does not. This study will help enable others in the robotics community to build more fluent and adaptable robots in the future. Tariq Iqbal, Samantha Rack, Laurel D. Riek |
IEEE Trans. Robotics | 1 |
| 2015 | Detecting and Synthesizing Synchronous Joint Action in Human-Robot TeamsabstractTo become capable teammates to people, robots need the ability to interpret human activities and appropriately adjust their actions in real time. The goal of our research is to build robots that can work fluently and contingently with human teams. To this end, we have designed novel nonlinear dynamical methods to automatically model and detect synchronous joint action (SJA) in human teams. We also have extended this work to enable robots to move jointly with human teammates in real time. In this paper, we describe our work to date, and discuss our future research plans to further explore this research space. The results of this work are expected to benefit researchers in social signal processing, human-machine interaction, and robotics. Tariq Iqbal, Laurel D. Riek |
ICMI | 1 |
| 2015 | Joint action perception to enable fluent human-robot teamworkabstractTo be effective team members, it is important for robots to understand the high-level behaviors of collocated humans. This is a challenging perceptual task when both the robots and people are in motion. In this paper, we describe an event-based model for multiple robots to automatically measure synchronous joint action of a group while both the robots and co-present humans are moving. We validated our model through an experiment where two people marched both synchronously and asynchronously, while being followed by two mobile robots. Our results suggest that our model accurately identifies synchronous motion, which can enable more adept human-robot collaboration. Tariq Iqbal, Michael J. Gonzales, Laurel D. Riek |
RO-MAN | 1 |