VLDB 2026 Research / reviewers in the wild / expert
Md. Mofijul Islam
dblp:168/9863 · also Md Mofijul Islam, Mofijul Islam
· DBLP profile ↗
12ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-4207-5863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Embodied Referring Expression Comprehension in Human-Robot InteractionabstractAs robots enter human workspaces, there is a crucial need for them to comprehend embodied human instructions, enabling intuitive and fluent human-robot interaction (HRI). However, accurate comprehension is challenging due to a lack of large-scale datasets that capture natural embodied interactions in diverse HRI settings. Existing datasets suffer from perspective bias, single-view data collection, inadequate coverage of nonverbal gestures, and a predominant focus on indoor environments. To address these issues, we present the Refer360 dataset, a large-scale dataset of embodied verbal and nonverbal interactions collected across diverse viewpoints in both indoor and outdoor settings. Additionally, we introduce MuRes, a multimodal guided residual module designed to improve embodied referring expression comprehension. MuRes acts as an information bottleneck, extracting salient modality-specific signals and reinforcing them into pre-trained representations to form complementary features for downstream tasks. We conduct extensive experiments on four datasets, including our Refer360 dataset, and demonstrate that current multimodal models fail to capture embodied interactions comprehensively; however, augmenting them with MuRes consistently improves performance. These findings establish Refer360 as a valuable benchmark and exhibit the potential of guided residual learning to advance embodied referring expression comprehension in robots operating within human environments. Md. Mofijul Islam, Alexi Gladstone, Sujan Sarker, Ganesh Nanduru, Md Fahim, Keyan Du, Aman Chadha, Tariq Iqbal |
HRI | 1 |
| 2024 | PoseTron: Enabling Close-Proximity Human-Robot Collaboration Through Multi-human Motion PredictionabstractAs robots enter human workspaces, there is a crucial need for robots to understand and predict human motion to achieve safe and fluent human-robot collaboration (HRC). However, accurate prediction is challenging due to a lack of large-scale datasets for close-proximity HRC and the absence of generalizable algorithms. To overcome these challenges, we present INTERACT, a comprehensive multimodal dataset covering 3-D Skeleton, RGB+D, gaze, and robot joint data for human-human and human-robot collaboration. Additionally, we introduce PoseTron, a novel transformer-based architecture to address the gap in learning algorithms. PoseTron introduces a conditional attention mechanism in the encoder enabling efficient weighing of motion information from all agents to incorporate team dynamics. The decoder features a novel multimodal attention mechanism, which weights representations from different modalities and the encoder outputs to predict future motion. We extensively evaluated PoseTron by comparing its performance on the INTERACT dataset against state-of-the-art algorithms. The results suggest that PoseTron outperformed all other methods across all the scenarios, attaining lowest prediction errors. Furthermore, we conducted a comprehensive ablation study, emphasizing the importance of design choices, pointing towards a promising direction for integrating motion prediction with robot perception in safe and effective HRC. Mohammad Samin Yasar, Md. Mofijul Islam, Tariq Iqbal |
HRI | 2 |
| 2024 | EQA-MX: Embodied Question Answering using Multimodal ExpressionabstractHumans predominantly use verbal utterances and nonverbal gestures (e.g., eye gaze and pointing gestures) in their natural interactions. For instance, pointing gestures and verbal information is often required to comprehend questions such as "what object is that?" Thus, this question-answering (QA) task involves complex reasoning of multimodal expressions (verbal utterances and nonverbal gestures). However, prior works have explored QA tasks in non-embodied settings, where questions solely contain verbal utterances from a single verbal and visual perspective. In this paper, we have introduced 8 novel embodied question answering (EQA) tasks to develop learning models to comprehend embodied questions with multimodal expressions. We have developed a novel large-scale dataset, EQA-MX, with over 8 million diverse embodied QA data samples involving multimodal expressions from multiple visual and verbal perspectives. To learn salient multimodal representations from discrete verbal embeddings and continuous wrapping of multiview visual representations, we propose a vector-quantization (VQ) based multimodal representation learning model, VQ-Fusion, for the EQA tasks. Our extensive experimental results suggest that VQ-Fusion can improve the performance of existing state-of-the-art visual-language models up to 13% across EQA tasks. Md. Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq Iqbal |
ICLR | 1 |
| 2024 | IMPRINT: Interactional Dynamics-aware Motion Prediction in Teams using Multimodal ContextabstractRobots are moving from working in isolation to working with humans as a part of human-robot teams. In such situations, they are expected to work with multiple humans and need to understand and predict the team members’ actions. To address this challenge, in this work, we introduce IMPRINT, a multi-agent motion prediction framework that models the interactional dynamics and incorporates the multimodal context (e.g., data from RGB and depth sensors and skeleton joint positions) to accurately predict the motion of all the agents in a team. In IMPRINT, we propose an Interaction module that can extract the intra-agent and inter-agent dynamics before fusing them to obtain the interactional dynamics. Furthermore, we propose a Multimodal Context module that incorporates multimodal context information to improve multi-agent motion prediction. We evaluated IMPRINT by comparing its performance on human-human and human-robot team scenarios against state-of-the-art methods. The results suggest that IMPRINT outperformed all other methods over all evaluated temporal horizons. Additionally, we provide an interpretation of how IMPRINT incorporates the multimodal context information from all the modalities during multi-agent motion prediction. The superior performance of IMPRINT provides a promising direction to integrate motion prediction with robot perception and enable safe and effective human-robot collaboration. Mohammad Samin Yasar, Md. Mofijul Islam, Tariq Iqbal |
ACM Trans. Hum. Robot Interact. | 2 |
| 2023 | PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal CuesabstractHumans naturally use referring expressions with verbal utterances and nonverbal gestures to refer to objects and events. As these referring expressions can be interpreted differently from the speaker's or the observer's perspective, people effectively decide on the perspective in comprehending the expressions. However, existing models do not explicitly learn perspective grounding, which often causes the models to perform poorly in understanding embodied referring expressions. To make it exacerbate, these models are often trained on datasets collected in non-embodied settings without nonverbal gestures and curated from an exocentric perspective. To address these issues, in this paper, we present a perspective-aware multitask learning model, called PATRON, for relation and object grounding tasks in embodied settings by utilizing verbal utterances and nonverbal cues. In PATRON, we have developed a guided fusion approach, where a perspective grounding task guides the relation and object grounding task. Through this approach, PATRON learns disentangled task-specific and task-guidance representations, where task-guidance representations guide the extraction of salient multimodal features to ground the relation and object accurately. Furthermore, we have curated a synthetic dataset of embodied referring expressions with multimodal cues, called CAESAR-PRO. The experimental results suggest that PATRON outperforms the evaluated state-of-the-art visual-language models. Additionally, the results indicate that learning to ground perspective helps machine learning models to improve the performance of the relation and object grounding task. Furthermore, the insights from the extensive experimental results and the proposed dataset will enable researchers to evaluate visual-language models' effectiveness in understanding referring expressions in other embodied settings. Md. Mofijul Islam, Alexi Gladstone, Tariq Iqbal |
AAAI | 1 |
| 2023 | Representation Learning in Deep RL via Discrete Information BottleneckabstractSeveral self-supervised representation learning methods have been proposed for reinforcement learning (RL) with rich observations. For real world applications of RL, recovering underlying latent states is crucial, particularly when sensory inputs can contain irrelevant and exogenous information. In this work, we study how information bottlenecjs can be used to construct latent states efficiently in the presence of task irrelevant information. We propose architectures that utilize variational and discrete information bottleneck, coined as RepDIB, to learn structured factorized representations. Exploiting the expressiveness bought by factorized representations, we introduce a simple, yet effective, bottleneck that can be integrated with any existing self supervised objective for RL. We demonstrate this across several online and offline RL benchmarks, along with a real robot arm task, where we find that compressed representations with RepDIB can lead to strong performance improvements, as the learnt bottlenecks can help predict only the relevant state, while ignoring irrelevant information. Riashat Islam, Hongyu Zang, Manan Tomar, Aniket Didolkar, Md. Mofijul Islam, Samin Yeasar Arnob, Tariq Iqbal, Xin Li 0033, Anirudh Goyal, Nicolas Heess, Alex Lamb |
AISTATS | 5 |
| 2023 | MAVEN: A Memory Augmented Recurrent Approach for Multimodal FusionabstractMultisensory systems provide complementary information that aids many machine learning approaches in perceiving the environment comprehensively. These systems consist of heterogeneous modalities, which have disparate characteristics and feature distributions. Thus, extracting, aligning, and fusing complementary representations from heterogeneous modalities (e.g., visual, skeleton, and physical sensors) remains challenging. To address these challenges, we have used the insights from several neuroscience studies of animal multisensory systems to develop MAVEN, a memory-augmented recurrent approach for multimodal fusion. MAVEN generates unimodal memory banks comprised of spatial-temporal features and uses our proposed recurrent representation alignment approach to align and refine unimodal representations iteratively. MAVEN then utilizes a multimodal variational attention-based fusion approach to produce a robust multimodal representation from the aligned unimodal features. Our extensive experimental evaluations on three multimodal datasets suggest that MAVEN outperforms state-of-the-art multimodal learning approaches in the challenging human activity recognition task across all evaluation conditions (cross-subject, leave-one-subject-out, and cross-session). Additionally, our extensive ablation studies suggest that MAVEN significantly outperforms the feed-forward fusion-based learning models$(p< 0.05)$. Finally, the robust performance of MAVEN in extracting complementary multimodal representation from occluded and noisy data suggests its applicability on real-world datasets. Md. Mofijul Islam, Mohammad Samin Yasar, Tariq Iqbal |
IEEE Trans. Multim. | 1 |
| 2022 | MuMu: Cooperative Multitask Learning-Based Guided Multimodal FusionabstractMultimodal sensors (visual, non-visual, and wearable) can provide complementary information to develop robust perception systems for recognizing activities accurately. However, it is challenging to extract robust multimodal representations due to the heterogeneous characteristics of data from multimodal sensors and disparate human activities, especially in the presence of noisy and misaligned sensor data. In this work, we propose a cooperative multitask learning-based guided multimodal fusion approach, MuMu, to extract robust multimodal representations for human activity recognition (HAR). MuMu employs an auxiliary task learning approach to extract features specific to each set of activities with shared characteristics (activity-group). MuMu then utilizes activity-group-specific features to direct our proposed Guided Multimodal Fusion Approach (GM-Fusion) for extracting complementary multimodal representations, designed as the target task. We evaluated MuMu by comparing its performance to state-of-the-art multimodal HAR approaches on three activity datasets. Our extensive experimental results suggest that MuMu outperforms all the evaluated approaches across all three datasets. Additionally, the ablation study suggests that MuMu significantly outperforms the baseline models (p Md. Mofijul Islam, Tariq Iqbal |
AAAI | 1 |
| 2022 | Who's Laughing NAO?: Examining Perceptions of Failure in a Humorous Robot PartnerabstractSocial robots are being deployed to interact with people in various scenarios, where they are expected to in-corporate human-like conversational strategies to achieve flu-ency in interactions. For example, current robots are designed to perform advanced communication strategies (i.e., personal anecdotes, explanations, and apologies) to recover from task failure. However, these tactics are not always sufficient for failure recovery as they can be lengthy and insufficient for encouraging future interactions. In human-human interactions, people often use humor as a low-risk and engaging method for managing failures. Thus, the successful execution of advanced, human-like humor could enable robots to recover from task failures more efficiently. In this paper, we present a human-robot interaction study exploring how a robot's utilization of various human-like humor types (i.e., affiliative, aggressive, self-enhancing, and self-defeating) are perceived by human teammate (n = 32) and an external observer of the interaction (n = 256). Additionally, we have explored the effects of performance, humor type, perspective, and previous experience with robots on the participants' perceptions of warmth, competence, and the robot as a teammate. Our results indicate that dyadic participants rated the successful robot to be more competent and a better teammate than the bystander participants. Additionally, the results indicate that participants with less experience with robots found the successful robot to be more competent than participants with high levels of experience. These findings will enable the human-robot interaction community to develop more engaging robots for fluent interactive experiences in the future. Haley N. Green, Md. Mofijul Islam, Shahira Ali, Tariq Iqbal |
HRI | 2 |
| 2022 | CAESAR: An Embodied Simulator for Generating Multimodal Referring Expression DatasetsabstractHumans naturally use verbal utterances and nonverbal gestures to refer to various objects (known as $\textit{referring expressions}$) in different interactional scenarios. As collecting real human interaction datasets are costly and laborious, synthetic datasets are often used to train models to unambiguously detect relationships among objects. However, existing synthetic data generation tools that provide referring expressions generally neglect nonverbal gestures. Additionally, while a few small-scale datasets contain multimodal cues (verbal and nonverbal), these datasets only capture the nonverbal gestures from an exo-centric perspective (observer). As models can use complementary information from multimodal cues to recognize referring expressions, generating multimodal data from multiple views can help to develop robust models. To address these critical issues, in this paper, we present a novel embodied simulator, CAESAR, to generate multimodal referring expressions containing both verbal utterances and nonverbal cues captured from multiple views. Using our simulator, we have generated two large-scale embodied referring expression datasets, which we have released publicly. We have conducted experimental analyses on embodied spatial relation grounding using various state-of-the-art baseline models. Our experimental results suggest that visual perspective affects the models' performance; and that nonverbal cues improve spatial relation grounding accuracy. Finally, we will release the simulator publicly to allow researchers to generate new embodied interaction datasets. Md. Mofijul Islam, Reza Mirzaiee, Alexi Gladstone, Haley N. Green, Tariq Iqbal |
NeurIPS | 1 |
| 2022 | Fusing Computer Vision and Wireless Signal for Accurate Sensor Localization in AR ViewabstractRecent years have seen increasing traction to enable new applications that can localize sensors on the screen of an Augmented Reality (AR) device (e.g. smartphone, tablet) so that sensors can be controlled more intuitively. Despite recent advances in this area, both wireless signal dependent and computer vision based localization solutions have seen a slow acceptance due to signal noise, multipath effect, and limited AR device-sensor interactivity. In this paper, we propose a novel solution to combine the complementary advantages of wireless signal based localization solution with the computer vision based solution to track IoT devices and sensors more accurately. Experimental result shows that our system can accurately track IoT devices with an average pixel error of 34 pixels in a 1024 × 768 pixels image, which is a 75.8% improvement from the state-of-the-art model. Md Fazlay Rabbi Masum Billah, Md. Mofijul Islam, Nurani Saoda, Tariq Iqbal, Bradford Campbell |
SenSys | 2 |
| 2020 | HAMLET: A Hierarchical Multimodal Attention-based Human Activity Recognition AlgorithmabstractTo fluently collaborate with people, robots need the ability to recognize human activities accurately. Although modern robots are equipped with various sensors, robust human activity recognition (HAR) still remains a challenging task for robots due to difficulties related to multimodal data fusion. To address these challenges, in this work, we introduce a deep neural network-based multimodal HAR algorithm, HAMLET. HAMLET incorporates a hierarchical architecture, where the lower layer encodes spatio-temporal features from unimodal data by adopting a multi-head self-attention mechanism. We develop a novel multimodal attention mechanism for disentangling and fusing the salient unimodal features to compute the multimodal features in the upper layer. Finally, multimodal features are used in a fully connect neural-network to recognize human activities. We evaluated our algorithm by comparing its performance to several state-of-the-art activity recognition algorithms on three human activity datasets. The results suggest that HAMLET outperformed all other evaluated baselines across all datasets and metrics tested, with the highest top-1 accuracy of 95.12% and 97.45% on the UTD-MHAD [1] and the UT-Kinect [2] datasets respectively, and F1-score of 81.52% on the UCSD-MIT [3] dataset. We further visualize the unimodal and multimodal attention maps, which provide us with a tool to interpret the impact of attention mechanisms concerning HAR. Md. Mofijul Islam, Tariq Iqbal |
IROS | 1 |