VLDB 2026 Research / reviewers in the wild / expert
Yahya Hmaiti
dblp:344/9267
· DBLP profile ↗
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0003-1052-1152ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From One World to Another: Interfaces for Efficiently Transitioning Between Virtual EnvironmentsabstractPersonal computers and handheld devices provide keyboard shortcuts and swipe gestures to enable users to efficiently switch between applications, whereas today’s virtual reality (VR) systems do not. In this work, we present an exploratory study on user interface aspects to support efficient switching between worlds in VR. We created eight interfaces that afford previewing and selecting from the available virtual worlds, including methods using portals and worlds-in-miniature (WiMs). To evaluate these methods, we conducted a controlled within-subjects empirical experiment (N=22) where participants frequently transitioned between six different environments to complete an object collection task. Our quantitative and qualitative results show that WiMs supported rapid acquisition of high-level spatial information while searching and were deemed most efficient by participants while portals provided fast pre-orientation. Finally, we present insights into the applicability, usability, and effectiveness of the VR world switching methods we explored, and provide recommendations for their application and future context/world switching techniques and interfaces. Matthew Gottsacker, Yahya Hmaiti, Mykola Maslych, Hiroshi Furuya, Jasmine DeGuzman, Gerd Bruder, Greg Welch, Joseph J. LaViola Jr. |
CHI | 2 |
| 2025 | The Fidelity-based Presence Scale (FPS): Modeling the Effects of Fidelity on Sense of PresenceabstractWithin the virtual reality (VR) research community, there have been several efforts to develop questionnaires with the aim of better understanding the sense of presence. Despite having numerous surveys, the community does not have a questionnaire that informs which components of a VR application contributed to the sense of presence. Furthermore, previous literature notes the absence of consensus on which questionnaire or questions should be used. Therefore, we conducted a Delphi study, engaging presence experts to establish a consensus on the most important presence questions and their respective verbiage. We then conducted a validation study with an exploratory factor analysis (EFA). The efforts between our two studies led to the creation of the Fidelity-based Presence Scale (FPS). With our consensus-driven approach and fidelity-based factoring, we hope the FPS will enable better communication within the research community and yield important future results regarding the relationship between VR system fidelity and presence. Jacob Belga, Richard Skarbez, Yahya Hmaiti, Eric J. Chen, Ryan P. McMahan, Joseph J. LaViola Jr. |
CHI | 3 |
| 2025 | All Languages Matter: Evaluating LMMs on Culturally Diverse 100 LanguagesabstractExisting Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model’s ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available at https://mbzuai-oryx.github.io/ALM-Bench/. Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana, Noor Ahsan, Nevasini Sasikumar, Omkar Thawakar, Henok Biadglign Ademtew, Yahya Hmaiti, Amandeep Kumar, Kartik Kuckreja, Mykola Maslych, Wafa Al Ghallabi, Mihail Minkov Mihaylov, Abdelrahman M. Shaker, Mike Zhang, Mahardika Krisna Ihsani, Amiel Esplana, Monil Gokani, Shachar Mirkin, Harsh Singh, Ashay Srivastava, Endre Hamerlik, Fathinah Asma Izzati, Fadillah A. Maani, Sebastian Cavada, Jenny Chim, Rohit Gupta 0012, Sanjay Manjunath, Kamila Zhumakhanova, Feno Heriniaina Rabevohitra, Azril Hafizi Amirudin, Muhammad Ridzuan, Daniya Najiha Abdul Kareem, Ketan More, Pramesh Shakya, Amirpouya Ghasemaghaei, Amirbek Djanibekov, Dilshod Azizov, Branislava Jankovic, Naman Bhatia, Alvaro Cabrera, Johan S. Obando-Ceron, Olympiah Otieno, Fabian Farestam, Muztoba Rabbani, Sanoojan Baliah, Santosh Sanjeev, Abduragim Shtanchaev, Maheen Fatima, Amrin Kareem, Toluwani Aremu, Nathan A. Z. Xavier, Amit Bhatkal, Hawau Olamide Toyin, Aman Chadha, Hisham Cholakkal, Rao Muhammad Anwer, Michael Felsberg, Jorma Laaksonen, Thamar Solorio, Monojit Choudhury, Ivan Laptev, Mubarak Shah, Salman Khan 0001, Fahad Shahbaz Khan |
CVPR | 8 |
| 2025 | A Culturally-diverse Multilingual Multimodal Video Benchmark & ModelabstractBhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin’ichi Satoh, Michael Felsberg, Mubarak Shah, Salman Khan, Fahad Shahbaz Khan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz 0001, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan More, Sanoojan Baliah, Hasindri Watawana, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber 0002, Hisham Cholakkal, Ivan Laptev, Shin'ichi Satoh 0001, Michael Felsberg, Mubarak Shah, Salman Khan 0001, Fahad Shahbaz Khan |
EMNLP | 7 |
| 2024 | Unlocking Understanding: An Investigation of Multimodal Communication in Virtual Reality CollaborationabstractCommunication in collaboration, especially synchronous, remote communication, is crucial to the success of task-specific goals. Insufficient or excessive forms of communication may lead to detrimental effects on task performance while increasing mental fatigue. However, identifying which combinations of communication modalities provide the most efficient transfer of information in collaborative settings will greatly improve collaboration. To investigate this, we developed a remote, synchronous, asymmetric VR collaborative assembly task application, where users play the role of either mentor or mentee, and were exposed to different combinations of three communication modalities: voice, gestures, and gaze. Through task-based experiments with 25 pairs of participants (50 individuals), we evaluated quantitative and qualitative data and found that gaze did not differ significantly from multiple combinations of communication modalities. Our qualitative results indicate that mentees experienced more difficulty and frustration in completing tasks than mentors, with both types of users preferring all three modalities to be present. Ryan Ghamandi, Ravi Kiran Kattoju, Yahya Hmaiti, Mykola Maslych, Eugene M. Taranta II, Ryan P. McMahan, Joseph J. LaViola Jr. |
CHI | 3 |
| 2024 | Towards Better Throwing: A Comparison of Performance and Preferences Across Point of Release Mechanics in Virtual RealityabstractAn underexplored interaction metaphor in virtual reality (VR) is throwing, with a considerable challenge in achieving accurate and natural results. We conducted an empirical investigation of participants’ performance in a VR throwing task, measuring their accuracy and preferences across Point of Release (PoR) mechanics (manual and automatic) with various input device categories (hand-held, on-body, external) and throwable object types. Participants were tasked with throwing a baseball, a bowling ball, and a football toward targets using 5 input configurations (2 manual and 3 automatic PoR). Results from 30 participants indicate that the overall highest accuracy was achieved with an automatic PoR configuration (on-body tracker). The post-study and VR survey results indicate that the majority of participants preferred a manual PoR configuration (hand-held VR controller-derived) for the throwing direction, throwing speed, and as being the closest to real-life throwing. Our findings are useful for VR researchers and developers who want to implement throwing as a technique in their applications. Amirpouya Ghasemaghaei, Mykola Maslych, Yahya Hmaiti, Esteban Segarra Martinez, Joseph J. LaViola Jr. |
Graphics Interface | 3 |
| 2024 | From Research to Practice: Survey and Taxonomy of Object Selection in Consumer VR ApplicationsabstractObject selection has been explored extensively in the VR research literature. However, the research is typically conducted in constrained experimental setups. It remains unclear whether the designed selection techniques fit the prevalent practical uses and whether the experimental tasks represent important challenges in real applications. To identify and help bridge these gaps, we surveyed current consumer VR applications, containing 206 popular VR game and 3D modeling applications. We extracted 1300+ selection scenarios based on video analyses of these applications and derived a taxonomy to understand common patterns on where and how selections occur. Our findings reveal significant gaps in selection tasks and techniques between research and consumer applications. We also present an interactive visualization tool to help researchers explore the VR object selection scenarios. Finally, we discuss how our work can help researchers and developers evaluate techniques in meaningful tasks and drive the design of techniques. Mykola Maslych, Difeng Yu, Amirpouya Ghasemaghaei, Yahya Hmaiti, Esteban Segarra Martinez, Dominic Simon, Eugene M. Taranta II, Joanna Bergström, Joseph J. LaViola Jr. |
ISMAR | 4 |
| 2024 | Visual Perceptual Confidence: Exploring Discrepancies Between Self-reported and Actual Distance Perception In Virtual RealityabstractVirtual Reality (VR) systems are widely used, and it is essential to know if spatial perception in virtual environments (VEs) is similar to reality. Research indicates that users tend to underestimate distances in VR. Prior work suggests that actual distance judgments in VR may not always match the users self-reported preference of where they think they most accurately estimated distances. However, no explicit investigation evaluated whether user preferences match actual performance in a spatial judgment task. We used blind walking to explore potential dissimilarities between actual distance estimates and user-selected preferences of visual complexities, VE conditions, and targets. Our findings show a gap between user preferences and actual performance when visual complexities were varied, which has implications for better visual perception understanding, VR applications design, and research in spatial perception, indicating the need to calibrate and align user preferences and true spatial perception abilities in VR. Yahya Hmaiti, Mykola Maslych, Amirpouya Ghasemaghaei, Ryan Ghamandi, Joseph J. LaViola Jr. |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | What And How Together: A Taxonomy On 30 Years Of Collaborative Human-Centered XR TasksabstractWe present a taxonomy of human-centered collaborative XR tasks. XR technologies have extended into the realm of collaboration, improving the quality and accessibility of teamwork. However, after a comprehensive assessment of the literature on the interaction between XR technologies and collaboration, no comprehensive method that emphasizes task actions and properties exists to classify collaborative tasks. Thus, our suggested taxonomy represents a classification system for collaborative tasks. After conducting a thorough literature review across different research venues, we conducted several exhaustive classification and review cycles for over 800 papers collected, which resulted in 148 papers retained to create the taxonomy. We dissected the actions and properties that the collaborative endeavors and tasks of these papers encompass as well as the types of categorizations and relations these papers illustrate. We expand on the design choices and usage of our taxonomy, followed by its limitations and future work. We built this taxonomy in order to reduce ambiguities and confusion regarding the design and comprehension of human-based collaborative tasks that use XR technology, which could prove useful in aiding the development and understanding of these tasks. Our taxonomy reveals a framework for understanding how collaborative tasks are designed and a systematic way of classifying different methods by which people can collaborate and interact in environments that involve XR, while still promoting efficient communication, teamwork, goal achievement and productivity. Ryan Ghamandi, Yahya Hmaiti, Tam T. Nguyen, Amirpouya Ghasemaghaei, Ravi Kiran Kattoju, Eugene M. Taranta II, Joseph J. LaViola Jr. |
ISMAR | 2 |
| 2023 | An Exploration of The Effects of Head-Centric Rest Frames On Egocentric Distance Judgments in VRabstractUsers tend to underestimate distances in virtual reality (VR), and several efforts have been directed toward finding the causes and developing tools that mitigate this phenomenon. One hypothesis that stands out in the field of spatial perception is the rest frame hypothesis (RFH), which states that visual frames of reference (RFs), defined as fixed reference points of view in a virtual environment (VE), contribute to minimizing sensory mismatch. RFs have been shown to promote better eye-gaze stability and focus, reduce VR sickness, and improve visual search, along with other benefits. However, their effect on distance perception in VEs has not been evaluated. In this paper, we use a blind walking task to explore the effect of three head-centric RFs (mesh mask, nose, and hat) on egocentric distance estimation. We found that at near and mid-field distances, certain RFs can improve the user’s distance estimation accuracy and reduce distance underestimation. These findings mean that the addition of head-centric RFs, a simple avatar augmentation method, can lead to meaningful improvements in distance judgments, user experience, and task performance in VR. Yahya Hmaiti, Mykola Maslych, Eugene M. Taranta II, Joseph J. LaViola Jr. |
ISMAR | 1 |
| 2023 | Toward Intuitive Acquisition of Occluded VR Objects Through an Interactive Disocclusion Mini-mapabstractStandard selection techniques such as ray casting fail when virtual objects are partially or fully occluded. In this paper, we present two novel approaches that combine cone-casting, world-in-miniature, and grasping metaphors to disocclude objects in the representation local to the user. Through a within-subject study where we compared 4 selection techniques across 3 levels of object occlusion, we found that our techniques outperformed an alternative one that also focuses on maintaining the spatial relationships between objects. We discuss application scenarios and future research directions for these types of selection techniques. Mykola Maslych, Yahya Hmaiti, Ryan Ghamandi, Paige Leber, Ravi Kiran Kattoju, Jacob Belga, Joseph J. LaViola Jr. |
VR | 2 |