VLDB 2026 Research / reviewers in the wild / expert
Mohammad Samin Yasar
dblp:234/8744
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0002-4684-2823ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Holistic Evaluation of Teleoperation Interfaces for Robotic ManipulationabstractRobot teleoperation has become increasingly crucial for extending human capabilities in inaccessible or hazardous environments and facilitating human–robot collaboration. While significant advancements have been made in teleoperation interfaces, the success of these systems critically depends on how effectively humans can interact with and control robotic systems across diverse manipulation tasks. However, existing research primarily evaluates interfaces within specific tasks or applications, lacking systematic assessment across different manipulation scenarios. This limitation leads to suboptimal interface selection that can compromise task efficiency in critical domains and impede the collection of high-quality demonstrations for robot learning. To address these gaps, we first introduce a novel two-axis Robotic Manipulation Task Taxonomy that systematically categorizes manipulation tasks based on their fundamental control requirements: Motion Type (translation-dominant vs. rotation-dominant) and Engagement Type (rigid vs. non-rigid object interactions). We conducted a comprehensive user study ( \(n=30\) ) evaluating three distinct teleoperation interfaces (Gamepad, 3D Mouse, and Virtual Reality (VR) Controller) based on this taxonomy. Our results indicate that the Gamepad and 3D mouse significantly outperformed the VR Controller in task completion time across all task categories. In contrast, the VR Controller showed higher first-attempt success rates but caused significantly greater cognitive load across all NASA Task Load Index (NASA-TLX) subscales compared to the Gamepad interface, and specifically higher mental demand, physical demand, and effort compared to the 3D mouse. Our results also indicate that although prior familiarity with an interface lowered perceived workload and enhanced perceived performance, it did not translate into improvements in actual task success rates or completion times. These insights provide valuable guidelines for optimizing teleoperation interfaces based on task requirements and highlight the importance of considering both cognitive demands and user experience in interface design. Shaid Hasan, Mohammad Samin Yasar, Tariq Iqbal |
ACM Trans. Hum. Robot Interact. | 2 |
| 2024 | PoseTron: Enabling Close-Proximity Human-Robot Collaboration Through Multi-human Motion PredictionabstractAs robots enter human workspaces, there is a crucial need for robots to understand and predict human motion to achieve safe and fluent human-robot collaboration (HRC). However, accurate prediction is challenging due to a lack of large-scale datasets for close-proximity HRC and the absence of generalizable algorithms. To overcome these challenges, we present INTERACT, a comprehensive multimodal dataset covering 3-D Skeleton, RGB+D, gaze, and robot joint data for human-human and human-robot collaboration. Additionally, we introduce PoseTron, a novel transformer-based architecture to address the gap in learning algorithms. PoseTron introduces a conditional attention mechanism in the encoder enabling efficient weighing of motion information from all agents to incorporate team dynamics. The decoder features a novel multimodal attention mechanism, which weights representations from different modalities and the encoder outputs to predict future motion. We extensively evaluated PoseTron by comparing its performance on the INTERACT dataset against state-of-the-art algorithms. The results suggest that PoseTron outperformed all other methods across all the scenarios, attaining lowest prediction errors. Furthermore, we conducted a comprehensive ablation study, emphasizing the importance of design choices, pointing towards a promising direction for integrating motion prediction with robot perception in safe and effective HRC. Mohammad Samin Yasar, Md. Mofijul Islam, Tariq Iqbal |
HRI | 1 |
| 2024 | M2RL: A Multimodal Multi-Interface Dataset for Robot Learning from Human DemonstrationsabstractImitation Learning, inspired by observational learning theory in cognitive psychology, is a promising approach for teaching robots to perform complex manipulation tasks. However, most imitation learning datasets exhibit biases by focusing on a single interface or modality when capturing human demonstrations. This limitation fails to fully capture the multimodal nature of how humans learn skills through demonstration. To bridge this gap, we introduce the M2RL dataset, a multimodal and multi-interface dataset collected from non-expert users across diverse manipulation tasks from four task categories using three distinct teleoperation interfaces. The M2RL dataset comprises RGB+D data from three camera perspectives (robot’s wrist and two exo-views), ego-view and gaze data from the human teleoperator’s perspective, and the robot’s proprioception data. Our extensive evaluation of state-of-the-art imitation learning algorithms on the M2RL dataset highlights the importance of multimodal and multi-interface data for learning robust policies for the robot. Additionally, the results indicate clear performance improvements when training on data from diverse interfaces and utilizing inputs from multiple camera streams. Our dataset and code are publicly available at: https://github.com/M2RL/m2rl-dataset. Shaid Hasan, Mohammad Samin Yasar, Tariq Iqbal |
ICMI | 2 |
| 2024 | IMPRINT: Interactional Dynamics-aware Motion Prediction in Teams using Multimodal ContextabstractRobots are moving from working in isolation to working with humans as a part of human-robot teams. In such situations, they are expected to work with multiple humans and need to understand and predict the team members’ actions. To address this challenge, in this work, we introduce IMPRINT, a multi-agent motion prediction framework that models the interactional dynamics and incorporates the multimodal context (e.g., data from RGB and depth sensors and skeleton joint positions) to accurately predict the motion of all the agents in a team. In IMPRINT, we propose an Interaction module that can extract the intra-agent and inter-agent dynamics before fusing them to obtain the interactional dynamics. Furthermore, we propose a Multimodal Context module that incorporates multimodal context information to improve multi-agent motion prediction. We evaluated IMPRINT by comparing its performance on human-human and human-robot team scenarios against state-of-the-art methods. The results suggest that IMPRINT outperformed all other methods over all evaluated temporal horizons. Additionally, we provide an interpretation of how IMPRINT incorporates the multimodal context information from all the modalities during multi-agent motion prediction. The superior performance of IMPRINT provides a promising direction to integrate motion prediction with robot perception and enable safe and effective human-robot collaboration. Mohammad Samin Yasar, Md. Mofijul Islam, Tariq Iqbal |
ACM Trans. Hum. Robot Interact. | 1 |
| 2023 | VADER: Vector-Quantized Generative Adversarial Network for Motion PredictionabstractHuman motion prediction is an essential component for enabling close-proximity human-robot collaboration. The task of accurately predicting human motion is non-trivial and is compounded by the variability of human motion and the presence of multiple humans in proximity. To address some of the open challenges in motion prediction, in this work, we propose VADER, a novel sequence learning algorithm that models past observed poses using a flexible discrete latent space. VADER introduces the concept of Vector Quantization for human motion prediction, enabling the learning of a discrete latent space without being restricted by any static prior. In addition, we propose a new objective function that uses the discriminator objective to penalize deviation of predicted motion from the ground-truth. Finally, to explicitly model interaction in multiple humans, we introduce a lightweight attention mechanism to condition per-agent prediction on the previous hidden states of all the agents. Our evaluation across three scenarios: single-agent, multi-agent, and human-robot collaboration shows that VADER outperformed all the state-of-the-art approaches, resulting in more feasible human poses that align better with the ground-truth. Finally, we conducted extensive ablation studies to emphasize the importance of the proposed modules. Mohammad Samin Yasar, Tariq Iqbal |
IROS | 1 |
| 2023 | MAVEN: A Memory Augmented Recurrent Approach for Multimodal FusionabstractMultisensory systems provide complementary information that aids many machine learning approaches in perceiving the environment comprehensively. These systems consist of heterogeneous modalities, which have disparate characteristics and feature distributions. Thus, extracting, aligning, and fusing complementary representations from heterogeneous modalities (e.g., visual, skeleton, and physical sensors) remains challenging. To address these challenges, we have used the insights from several neuroscience studies of animal multisensory systems to develop MAVEN, a memory-augmented recurrent approach for multimodal fusion. MAVEN generates unimodal memory banks comprised of spatial-temporal features and uses our proposed recurrent representation alignment approach to align and refine unimodal representations iteratively. MAVEN then utilizes a multimodal variational attention-based fusion approach to produce a robust multimodal representation from the aligned unimodal features. Our extensive experimental evaluations on three multimodal datasets suggest that MAVEN outperforms state-of-the-art multimodal learning approaches in the challenging human activity recognition task across all evaluation conditions (cross-subject, leave-one-subject-out, and cross-session). Additionally, our extensive ablation studies suggest that MAVEN significantly outperforms the feed-forward fusion-based learning models$(p< 0.05)$. Finally, the robust performance of MAVEN in extracting complementary multimodal representation from occluded and noisy data suggests its applicability on real-world datasets. Md. Mofijul Islam, Mohammad Samin Yasar, Tariq Iqbal |
IEEE Trans. Multim. | 2 |
| 2022 | Robots That Can Anticipate and Learn in Human-Robot TeamsabstractRobots are moving from working in isolated cham-bers to working in close-proximity with human collaborator(s) as part of human-robot teams. In such situations, robots are increasingly expected to work with multiple humans and ef-fectively model both human-human and human-robot dynamics before taking timely actions. Working toward this goal, we have proposed new algorithms that model human intent and motion while being interpretable and scalable to multiple humans. Our current work builds upon these algorithms to 1) obtain a more holistic representation of the environment and 2) interleave robot perception and control. Our proposed algorithms have attained state-of-the-art performances over various benchmarks and learning scenarios. As part of future work, we aim to enhance our learning algorithms with the capability of acquiring knowledge continually, without overwriting past information. Mohammad Samin Yasar, Tariq Iqbal |
HRI | 1 |
| 2020 | Real-Time Context-Aware Detection of Unsafe Events in Robot-Assisted SurgeryabstractCyber-physical systems for robotic surgery have enabled minimally invasive procedures with increased precision and shorter hospitalization. However, with increasing complexity and connectivity of software and major involvement of human operators in the supervision of surgical robots, there remain significant challenges in ensuring patient safety. This paper presents a safety monitoring system that, given the knowledge of the surgical task being performed by the surgeon, can detect safety-critical events in real-time. Our approach integrates a surgical gesture classifier that infers the operational context from the time-series kinematics data of the robot with a library of erroneous gesture classifiers that given a surgical gesture can detect unsafe events. Our experiments using data from two surgical platforms show that the proposed system can detect unsafe events caused by accidental or malicious faults within an average reaction time window of 1,693 milliseconds and F1 score of 0.88 and human errors within an average reaction time window of 57 milliseconds and F1 score of 0.76. Mohammad Samin Yasar, Homa Alemzadeh |
DSN | 1 |