Mykola Maslych

dblp:264/7419 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0001-7037-3513ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 From One World to Another: Interfaces for Efficiently Transitioning Between Virtual Environments
abstract
Personal computers and handheld devices provide keyboard shortcuts and swipe gestures to enable users to efficiently switch between applications, whereas today’s virtual reality (VR) systems do not. In this work, we present an exploratory study on user interface aspects to support efficient switching between worlds in VR. We created eight interfaces that afford previewing and selecting from the available virtual worlds, including methods using portals and worlds-in-miniature (WiMs). To evaluate these methods, we conducted a controlled within-subjects empirical experiment (N=22) where participants frequently transitioned between six different environments to complete an object collection task. Our quantitative and qualitative results show that WiMs supported rapid acquisition of high-level spatial information while searching and were deemed most efficient by participants while portals provided fast pre-orientation. Finally, we present insights into the applicability, usability, and effectiveness of the VR world switching methods we explored, and provide recommendations for their application and future context/world switching techniques and interfaces.
Matthew Gottsacker, Yahya Hmaiti, Mykola Maslych, Hiroshi Furuya, Jasmine DeGuzman, Gerd Bruder, Greg Welch, Joseph J. LaViola Jr.
CHI3
2025 All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages
abstract
Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model’s ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available at https://mbzuai-oryx.github.io/ALM-Bench/.
Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana, Noor Ahsan, Nevasini Sasikumar, Omkar Thawakar, Henok Biadglign Ademtew, Yahya Hmaiti, Amandeep Kumar, Kartik Kuckreja, Mykola Maslych, Wafa Al Ghallabi, Mihail Minkov Mihaylov, Abdelrahman M. Shaker, Mike Zhang, Mahardika Krisna Ihsani, Amiel Esplana, Monil Gokani, Shachar Mirkin, Harsh Singh, Ashay Srivastava, Endre Hamerlik, Fathinah Asma Izzati, Fadillah A. Maani, Sebastian Cavada, Jenny Chim, Rohit Gupta 0012, Sanjay Manjunath, Kamila Zhumakhanova, Feno Heriniaina Rabevohitra, Azril Hafizi Amirudin, Muhammad Ridzuan, Daniya Najiha Abdul Kareem, Ketan More, Pramesh Shakya, Amirpouya Ghasemaghaei, Amirbek Djanibekov, Dilshod Azizov, Branislava Jankovic, Naman Bhatia, Alvaro Cabrera, Johan S. Obando-Ceron, Olympiah Otieno, Fabian Farestam, Muztoba Rabbani, Sanoojan Baliah, Santosh Sanjeev, Abduragim Shtanchaev, Maheen Fatima, Amrin Kareem, Toluwani Aremu, Nathan A. Z. Xavier, Amit Bhatkal, Hawau Olamide Toyin, Aman Chadha, Hisham Cholakkal, Rao Muhammad Anwer, Michael Felsberg, Jorma Laaksonen, Thamar Solorio, Monojit Choudhury, Ivan Laptev, Mubarak Shah, Salman Khan 0001, Fahad Shahbaz Khan
CVPR11
2025 A Culturally-diverse Multilingual Multimodal Video Benchmark & Model
abstract
Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan Pravin More, Sanoojan Baliah, Hasindri Watawana, Yuhao Li, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber, Hisham Cholakkal, Ivan Laptev, Shin’ichi Satoh, Michael Felsberg, Mubarak Shah, Salman Khan, Fahad Shahbaz Khan. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Bhuiyan Sanjid Shafique, Ashmal Vayani, Muhammad Maaz 0001, Hanoona Abdul Rasheed, Dinura Dissanayake, Mohammed Irfan Kurpath, Yahya Hmaiti, Go Inoue, Jean Lahoud, Md. Safirur Rashid, Shadid Intisar Quasem, Maheen Fatima, Franco Vidal, Mykola Maslych, Ketan More, Sanoojan Baliah, Hasindri Watawana, Fabian Farestam, Leon Schaller, Roman Tymtsiv, Simon Weber 0002, Hisham Cholakkal, Ivan Laptev, Shin'ichi Satoh 0001, Michael Felsberg, Mubarak Shah, Salman Khan 0001, Fahad Shahbaz Khan
EMNLP14
2024 Unlocking Understanding: An Investigation of Multimodal Communication in Virtual Reality Collaboration
abstract
Communication in collaboration, especially synchronous, remote communication, is crucial to the success of task-specific goals. Insufficient or excessive forms of communication may lead to detrimental effects on task performance while increasing mental fatigue. However, identifying which combinations of communication modalities provide the most efficient transfer of information in collaborative settings will greatly improve collaboration. To investigate this, we developed a remote, synchronous, asymmetric VR collaborative assembly task application, where users play the role of either mentor or mentee, and were exposed to different combinations of three communication modalities: voice, gestures, and gaze. Through task-based experiments with 25 pairs of participants (50 individuals), we evaluated quantitative and qualitative data and found that gaze did not differ significantly from multiple combinations of communication modalities. Our qualitative results indicate that mentees experienced more difficulty and frustration in completing tasks than mentors, with both types of users preferring all three modalities to be present.
Ryan Ghamandi, Ravi Kiran Kattoju, Yahya Hmaiti, Mykola Maslych, Eugene M. Taranta II, Ryan P. McMahan, Joseph J. LaViola Jr.
CHI4
2024 Towards Better Throwing: A Comparison of Performance and Preferences Across Point of Release Mechanics in Virtual Reality
abstract
An underexplored interaction metaphor in virtual reality (VR) is throwing, with a considerable challenge in achieving accurate and natural results. We conducted an empirical investigation of participants’ performance in a VR throwing task, measuring their accuracy and preferences across Point of Release (PoR) mechanics (manual and automatic) with various input device categories (hand-held, on-body, external) and throwable object types. Participants were tasked with throwing a baseball, a bowling ball, and a football toward targets using 5 input configurations (2 manual and 3 automatic PoR). Results from 30 participants indicate that the overall highest accuracy was achieved with an automatic PoR configuration (on-body tracker). The post-study and VR survey results indicate that the majority of participants preferred a manual PoR configuration (hand-held VR controller-derived) for the throwing direction, throwing speed, and as being the closest to real-life throwing. Our findings are useful for VR researchers and developers who want to implement throwing as a technique in their applications.
Amirpouya Ghasemaghaei, Mykola Maslych, Yahya Hmaiti, Esteban Segarra Martinez, Joseph J. LaViola Jr.
Graphics Interface2
2024 From Research to Practice: Survey and Taxonomy of Object Selection in Consumer VR Applications
abstract
Object selection has been explored extensively in the VR research literature. However, the research is typically conducted in constrained experimental setups. It remains unclear whether the designed selection techniques fit the prevalent practical uses and whether the experimental tasks represent important challenges in real applications. To identify and help bridge these gaps, we surveyed current consumer VR applications, containing 206 popular VR game and 3D modeling applications. We extracted 1300+ selection scenarios based on video analyses of these applications and derived a taxonomy to understand common patterns on where and how selections occur. Our findings reveal significant gaps in selection tasks and techniques between research and consumer applications. We also present an interactive visualization tool to help researchers explore the VR object selection scenarios. Finally, we discuss how our work can help researchers and developers evaluate techniques in meaningful tasks and drive the design of techniques.
Mykola Maslych, Difeng Yu, Amirpouya Ghasemaghaei, Yahya Hmaiti, Esteban Segarra Martinez, Dominic Simon, Eugene M. Taranta II, Joanna Bergström, Joseph J. LaViola Jr.
ISMAR1
2024 Visual Perceptual Confidence: Exploring Discrepancies Between Self-reported and Actual Distance Perception In Virtual Reality
abstract
Virtual Reality (VR) systems are widely used, and it is essential to know if spatial perception in virtual environments (VEs) is similar to reality. Research indicates that users tend to underestimate distances in VR. Prior work suggests that actual distance judgments in VR may not always match the users self-reported preference of where they think they most accurately estimated distances. However, no explicit investigation evaluated whether user preferences match actual performance in a spatial judgment task. We used blind walking to explore potential dissimilarities between actual distance estimates and user-selected preferences of visual complexities, VE conditions, and targets. Our findings show a gap between user preferences and actual performance when visual complexities were varied, which has implications for better visual perception understanding, VR applications design, and research in spatial perception, indicating the need to calibrate and align user preferences and true spatial perception abilities in VR.
Yahya Hmaiti, Mykola Maslych, Amirpouya Ghasemaghaei, Ryan Ghamandi, Joseph J. LaViola Jr.
IEEE Trans. Vis. Comput. Graph.2
2023 Effective 2D Stroke-based Gesture Augmentation for RNNs
abstract
Recurrent neural networks (RNN) require large training datasets from which they learn new class models. This limitation prohibits their use in custom gesture applications where only one or two end user samples are given per gesture class. One common way to enhance sparse datasets is to use data augmentation to synthesize new samples. Although there are numerous known techniques, they are often treated as standalone approaches when in reality they are often complementary. We show that by intelligently chaining augmentation techniques together that simulate different gesture production variability types, such as those affecting the temporal and spatial qualities of a gesture, we can significantly increase RNN accuracy without sacrificing training time. Through experimentation on four public stroke-based 2D gesture datasets, we show that RNNs trained with our data augmentation chaining technique achieves state-of-the-art recognition accuracy in both writer-dependent and writer-independent test scenarios.
Mykola Maslych, Eugene M. Taranta II, Mostafa Aldilati, Joseph J. LaViola Jr.
CHI1
2023 An Exploration of The Effects of Head-Centric Rest Frames On Egocentric Distance Judgments in VR
abstract
Users tend to underestimate distances in virtual reality (VR), and several efforts have been directed toward finding the causes and developing tools that mitigate this phenomenon. One hypothesis that stands out in the field of spatial perception is the rest frame hypothesis (RFH), which states that visual frames of reference (RFs), defined as fixed reference points of view in a virtual environment (VE), contribute to minimizing sensory mismatch. RFs have been shown to promote better eye-gaze stability and focus, reduce VR sickness, and improve visual search, along with other benefits. However, their effect on distance perception in VEs has not been evaluated. In this paper, we use a blind walking task to explore the effect of three head-centric RFs (mesh mask, nose, and hat) on egocentric distance estimation. We found that at near and mid-field distances, certain RFs can improve the user’s distance estimation accuracy and reduce distance underestimation. These findings mean that the addition of head-centric RFs, a simple avatar augmentation method, can lead to meaningful improvements in distance judgments, user experience, and task performance in VR.
Yahya Hmaiti, Mykola Maslych, Eugene M. Taranta II, Joseph J. LaViola Jr.
ISMAR2
2023 Toward Intuitive Acquisition of Occluded VR Objects Through an Interactive Disocclusion Mini-map
abstract
Standard selection techniques such as ray casting fail when virtual objects are partially or fully occluded. In this paper, we present two novel approaches that combine cone-casting, world-in-miniature, and grasping metaphors to disocclude objects in the representation local to the user. Through a within-subject study where we compared 4 selection techniques across 3 levels of object occlusion, we found that our techniques outperformed an alternative one that also focuses on maintaining the spatial relationships between objects. We discuss application scenarios and future research directions for these types of selection techniques.
Mykola Maslych, Yahya Hmaiti, Ryan Ghamandi, Paige Leber, Ravi Kiran Kattoju, Jacob Belga, Joseph J. LaViola Jr.
VR1
2022 The Voight-Kampff Machine for Automatic Custom Gesture Rejection Threshold Selection
abstract
Gesture recognition systems using nearest neighbor pattern matching are able to distinguish gesture from non-gesture actions by rejecting input whose recognition scores are poor. However, in the context of gesture customization, where training data is sparse, learning a tight rejection threshold that maximizes accuracy in the presence of continuous high activity (HA) data is a challenging problem. To this end, we present the Voight-Kampff Machine (VKM), a novel approach for rejection threshold selection. VKM uses new synthetic data techniques to select an initial threshold that the system thereafter adjusts based on the training set size and expected gesture production variability. We pair VKM with a state-of-the-art custom gesture segmenter and recognizer to evaluate our system across several HA datasets, where gestures are interleaved with non-gesture actions. Compared to alternative rejection threshold selection techniques, we show that our approach is the only one that consistently achieves high performance.
Eugene M. Taranta II, Mykola Maslych, Ryan Ghamandi, Joseph J. LaViola Jr.
CHI2
2021 Machete: Easy, Efficient, and Precise Continuous Custom Gesture Segmentation
abstract
We present Machete, a straightforward segmenter one can use to isolate custom gestures in continuous input. Machete uses traditional continuous dynamic programming with a novel dissimilarity measure to align incoming data with gesture class templates in real time. Advantages of Machete over alternative techniques is that our segmenter is computationally efficient, accurate, device-agnostic, and works with a single training sample. We demonstrate Machete’s effectiveness through an extensive evaluation using four new high-activity datasets that combine puppeteering, direct manipulation, and gestures. We find that Machete outperforms three alternative techniques in segmentation accuracy and latency, making Machete the most performant segmenter. We further show that when combined with a custom gesture recognizer, Machete is the only option that achieves both high recognition accuracy and low latency in a video game application.
Eugene M. Taranta II, Corey Pittman, Mehran Maghoumi, Mykola Maslych, Yasmine M. Moolenaar, Joseph J. LaViola Jr.
ACM Trans. Comput. Hum. Interact.4
2020 Moving Toward an Ecologically Valid Data Collection Protocol for 2D Gestures In Video Games
abstract
Those who design gesture recognizers and user interfaces often use data collection applications that enable users to comfortably produce gesture training samples. In contrast, games present unique contexts that impact cognitive load and have the potential to elicit rapid gesticulations as players react to dynamic conditions, which can result in high gesture form variability. However, the extent to which these gestures differ is presently unknown. To this end, we developed two games with unique mechanics, Follow the Leader (FTL) and Sleepy Town, as well as a standard data collection application. We collected gesture samples from 18 participants across all conditions for gestures of varying complexity, and through an analysis using relative, global, and distribution coverage measures, we confirm significant differences between conditions. We discuss the implications of our findings, and show that our FTL design is closer to being an ecologically valid data collection protocol with low implementation complexity.
Eugene M. Taranta II, Corey Pittman, Jack P. Oakley, Mykola Maslych, Mehran Maghoumi, Joseph J. LaViola Jr.
CHI4