EDBT 2026 Demo / reviewers in the wild / expert
Mubbasir Kapadia
dblp:08/4943
· DBLP profile ↗
121ranked-venue papers
13as first author
41since 2021 · last 2026
0000-0002-3501-0028ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 98 · 12 first-author · 29 since 2021Artificial intelligence and machine learning · 66 · 5 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 25 · 7 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Systems, architecture and hardware · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FCC: Fully Connected Correlation for One-Shot SegmentationabstractOne-shot segmentation (OSS) aims to segment the target object in a query image using only one set of support image and mask. Therefore, having strong prior information for the target object using the support set is essential to guide the initial training of OSS, which leads to the success of one-shot segmentation in challenging cases, such as when the target object shows considerable variation in appearance, texture, or scale across the support and query images. To enrich this prior knowledge, we introduce FCC (Fully Connected Correlation) which integrates pixel-level correlations between support and query features, capturing associations that reveal target-specific patterns and correspondences in both same-layers and cross-layers. FCC captures previously inaccessible target information, effectively addressing the limitations of support mask. Our approach consistently demonstrates state-of-the-art performance in the PASCAL, COCO, and domain shift tests, while also notably accelerating model convergence. We conducted an ablation study and cross-layer correlation analysis to validate FCC’s core methodology. These findings reveal the effectiveness of FCC in enhancing prior information and overall model performance for OSS1. Seonghyeon Moon, Haein Kong, Muhammad Haris Khan, Mubbasir Kapadia, Yuewei Lin |
WACV | 4 |
| 2026 | Large Sign Language Models: Toward 3D American Sign Language TranslationabstractWe present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which can benefit hearing-impaired individuals’ virtual communication. Unlike existing sign language recognition methods that rely on 2D video, our approach directly utilizes 3D sign language data to capture rich spatial, gestural, and depth information in 3D scenes. This enables more accurate and resilient translation, enhancing digital communication accessibility for the hearing-impaired community. Beyond the task of ASL translation, our work explores the integration of complex, embodied multimodal languages into the processing capabilities of LLMs, moving beyond purely text-based inputs to broaden their understanding of human communication. We investigate both direct translation from 3D gesture features to text and an instruction-guided setting where translations can be modulated by external prompts, offering greater flexibility. This work provides a foundational step toward inclusive, multimodal intelligent systems capable of understanding diverse forms of language. Xiaoxiao He, Di Liu 0003, Zhaoyang Xia, Chaowei Tan, Vivian Li, Bo Liu 0005, Dimitris N. Metaxas, Mubbasir Kapadia |
WACV | 10 |
| 2025 | On the Separability of Human Navigational Behaviors in Virtual Reality
Samuel S. Sohn, Serena DeStefani, Mathew Schwartz, Jacob Feldman, Mubbasir Kapadia, Karin Stromswold |
CogSci | 5 |
| 2025 | Cardiverse: Harnessing LLMs for Novel Card Game PrototypingabstractThe prototyping of computer games, particularly card games, requires extensive human effort in creative ideation and gameplay evaluation.Recent advances in Large Language Models (LLMs) offer opportunities to automate and streamline these processes.However, it remains challenging for LLMs to design novel game mechanics beyond existing databases, generate consistent gameplay environments, and develop scalable gameplay AI for large-scale evaluations.This paper addresses these challenges by introducing a comprehensive automated card game prototyping framework.The approach highlights a graph-based indexing method for generating novel game variations, an LLM-driven system for consistent game code generation validated by gameplay records, and a gameplay AI constructing method that uses an ensemble of LLM-generated heuristic functions optimized through self-play.These contributions aim to accelerate card game prototyping, reduce human labor, and lower barriers to entry for game developers. Danrui Li, Samuel S. Sohn, Kaidong Hu, Muhammad Usman 0010, Mubbasir Kapadia |
EMNLP | 6 |
| 2025 | Less is More: Improving Motion Diffusion Models with Sparse KeyframesabstractRecent advances in motion diffusion models have led to remarkable progress in diverse motion generation tasks, including text-to-motion synthesis. However, existing approaches represent motions as dense frame sequences, requiring the model to process redundant or less informative frames. The processing of dense animation frames imposes significant training complexity, especially when learning intricate distributions of large motion datasets even with modern neural architectures. This severely limits the performance of generative motion models for downstream tasks. Inspired by professional animators who mainly focus on sparse keyframes, we propose a novel diffusion framework explicitly designed around sparse and geometrically meaningful keyframes. Our method reduces computation by masking non-keyframes and efficiently interpolating missing frames. We dynamically refine the keyframe mask during inference to prioritize informative frames in later diffusion steps. Extensive experiments show that our approach consistently outperforms state-of-the-art methods in text alignment and motion realism, while also effectively maintaining high performance at significantly fewer diffusion steps. We further validate the robustness of our framework by using it as a generative prior and adapting it to different downstream tasks. Jinseok Bae, Inwoo Hwang, Young Yoon Lee, Yizhak Ben-Shabat, Young Min Kim 0001, Mubbasir Kapadia |
ICCV | 8 |
| 2025 | StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross FusionabstractWe present StyleMotif, a novel Stylized Motion Latent Diffusion model, generating motion conditioned on both content and style from multiple modalities. Unlike existing approaches that either focus on generating diverse motion content or transferring style from sequences, StyleMotif seamlessly synthesizes motion across a wide range of content while incorporating stylistic cues from multi-modal inputs, including motion, text, image, video, and audio. To achieve this, we introduce a style-content cross fusion mechanism and align a style encoder with a pre-trained multi-modal model, ensuring that the generated motion accurately captures the reference style while preserving realism. Extensive experiments demonstrate that our framework surpasses existing methods in stylized motion generation and exhibits emergent capabilities for multi-modal motion stylization, enabling more nuanced motion synthesis. Source code and pre-trained models will be released upon acceptance. Project Page: https://stylemotif.github.io Yizhak Ben-Shabat, Young Yoon Lee, Victor Zordan, Mubbasir Kapadia |
ICCV | 6 |
| 2025 | Adversarial Reinforcement Learning for Enhanced Decision-Making of Evacuation Guidance Robots in Intelligent Fire ScenariosabstractIn the context of rapid urbanization, traditional manual guidance and static evacuation signs are increasingly inadequate for addressing complex and dynamic emergencies. This study proposes an innovative emergency evacuation framework that optimizes the crowd evacuation by integrating multiagent reinforcement learning (MARL) with adversarial reinforcement learning (ARL). The developed simulation environment models realistic human behavior in complex buildings and incorporates robotic navigation and intelligent path planning. A novel simulated human behavior model was integrated, capable of complex human–robot interaction, independent escape route searching, and exhibiting herd mentality and memory mechanisms. We also proposed a multiagent framework that combines MARL and ARL to enhance overall evacuation efficiency and robustness. Additionally, we developed a new ARL evaluation framework that provides a novel method for quantifying agents’ performance. Various experiments of differing difficulty levels were conducted, and the results demonstrate that the proposed framework exhibits advantages in emergency evacuation scenarios. Specifically, our ARLR approach increased survival rates by 1.8% points in low-difficulty evacuation tasks compared to the RLR approach using only MARL algorithms. In high-difficulty evacuation tasks, the ARLR approach raised survival rates from 46.7% without robots to 64.4%, exceeding the RLR approach by 1.7% points. This study aims to enhance the efficiency and safety of human–robot collaborative fire evacuations and provides theoretical support for evaluating and improving the performance and robustness of ARL agents. Hantao Zhao, Tianxing Ma, Xiaomeng Shi, Mubbasir Kapadia, Tyler Thrash, Christoph Hölscher, Jinyuan Jia 0002, Bo Liu 0004, Jiuxin Cao |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2024 | Learning from Synthetic Human Group ActivitiesabstractThe study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation, we introduce M3 Act, a synthetic data generator for multi-view multi-group multi-person human atomic actions and group activities. Powered by Unity Engine, M3 Act features mul-tiple semantic groups, highly diverse and photorealistic images, and a comprehensive set of annotations, which facilitates the learning of human-centered tasks across single-person, multi-person, and multi-group conditions. We demonstrate the advantages of M3 Act across three core experiments. The results suggest our synthetic dataset can significantly improve the performance of several downstream methods and replace real-world datasets to reduce cost. Notably, M3 Act improves the state-of-the-art MOTRv2 on DanceTrack dataset, leading to a hop on the leaderboard from 10thto 2ndplace. Moreover, M3 Act opens new research for controllable 3D group activity generation. We define multiple metrics and propose a competitive baseline for the novel task. Our code and data are available at our project page: http://cjerry1243.github.io/M3Act. Che-Jui Chang, Danrui Li, Deep Patel, Parth Goel, Honglu Zhou, Seonghyeon Moon, Samuel S. Sohn, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
CVPR | 10 |
| 2024 | An Intrinsic Vector Heat NetworkabstractVector fields are widely used to represent and model flows for many science and engineering applications. This paper introduces a novel neural network architecture for learning tangent vector fields that are intrinsically defined on manifold surfaces embedded in 3D. Previous approaches to learning vector fields on surfaces treat vectors as multi-dimensional scalar fields, using traditional scalar-valued architectures to process channels individually, thus fail to preserve fundamental intrinsic properties of the vector field. The core idea of this work is to introduce a trainable vector heat diffusion module to spatially propagate vector-valued feature data across the surface, which we incorporate into our proposed architecture that consists of vector-valued neurons. Our architecture is invariant to rigid motion of the input, isometric deformation, and choice of local tangent bases, and is robust to discretizations of the surface. We evaluate our Vector Heat Network on triangle meshes, and empirically validate its invariant properties. We also demonstrate the effectiveness of our method on the useful industrial application of quadrilateral mesh generation. Alexander Gao, Maurice Chu, Mubbasir Kapadia, Ming C. Lin, Hsueh-Ti Derek Liu |
ICML | 3 |
| 2024 | TrajDiffuse: A Conditional Diffusion Model for Environment-Aware Trajectory Prediction
Tony Qingze Liu, Danrui Li, Samuel S. Sohn, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001 |
ICPR (29) | 5 |
| 2024 | From Words to Worlds: Transforming One-line Prompts into Multi-modal Digital Stories with LLM AgentsabstractDigital storytelling, essential in entertainment, education, and marketing, faces challenges in generation efficiency. The StoryAgent framework, introduced in this paper, utilizes Large Language Models and generative tools to automate and refine digital storytelling. Employing a top-down story drafting and bottom-up asset generation approach, StoryAgent tackles key issues such as manual intervention, interactive scene orchestration, and narrative consistency. This framework enables efficient production of interactive and consistent digital storytellings across multiple modalities, democratizing content creation and enhancing engagement. Danrui Li, Samuel S. Sohn, Che-Jui Chang, Mubbasir Kapadia |
MIG | 5 |
| 2023 | Procedure-Aware Pretraining for Instructional Video UnderstandingabstractOur goal is to learn a video representation that is useful for downstream procedure understanding tasks in instructional videos. Due to the small amount of available annotations, a key challenge in procedure understanding is to be able to extract from unlabeled videos the procedural knowledge such as the identity of the task (e.g., ‘make latte’), its steps (e.g., ‘pour milk’), or the potential next steps given partial progress in its execution. Our main insight is that instructional videos depict sequences of steps that repeat between instances of the same or different tasks, and that this structure can be well represented by a Procedural Knowledge Graph (PKG), where nodes are discrete steps and edges connect steps that occur sequentially in the instructional activities. This graph can then be used to generate pseudo labels to train a video representation that encodes the procedural knowledge in a more accessible form to generalize to multiple procedure understanding tasks. We build a PKG by combining information from a text-based procedural knowledge database and an unlabeled instructional video corpus and then use it to generate training pseudo labels with four novel pre-training objectives. We call this PKG-based pre-training procedure and the resulting model Paprika, Procedure-Aware PRe-training for Instructional Knowledge Acquisition. We evaluate Paprika on COIN and CrossTask for procedure understanding tasks such as task recognition, step recognition, and step forecasting. Paprika yields a video representation that improves over the state of the art: up to 11.23% gains in accuracy in 12 evaluation settings. Implementation is available at https://github.com/salesforce/paprika. Honglu Zhou, Roberto Martin Martin, Mubbasir Kapadia, Silvio Savarese, Juan Carlos Niebles |
CVPR | 3 |
| 2023 | MSI: Maximize Support-Set Information for Few-Shot SegmentationabstractFSS (Few-shot segmentation) aims to segment a target class using a small number of labeled images (support set). To extract information relevant to the target class, a dominant approach in best performing FSS methods removes background features using a support mask. We observe that this feature excision through a limiting support mask introduces an information bottleneck in several challenging FSS cases, e.g., for small targets and/or inaccurate target boundaries. To this end, we present a novel method (MSI), which maximizes the support-set information by exploiting two complementary sources of features to generate super correlation maps. We validate the effectiveness of our approach by instantiating it into three recent and strong FSS methods. Experimental results on several publicly available FSS benchmarks show that our proposed method consistently improves performance by visible margins and leads to faster convergence. Our code and trained models are available at: https://github.com/moonsh/MSI-Maximize-Support-Set-Information Seonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon, Vladimir Pavlovic 0001, Muhammad Haris Khan, Mubbasir Kapadia |
ICCV | 7 |
| 2023 | Harnessing Neighborhood Modeling and Asymmetry Preservation for Digraph Representation LearningabstractDigraph Representation Learning aims to learn representations for directed homogeneous graphs (digraphs). Prior work is largely constrained or has poor generalizability across tasks. Most Graph Neural Networks exhibit poor performance on digraphs due to the neglect of modeling neighborhoods and preserving asymmetry. In this paper, we address these notable challenges by leveraging hyperbolic collaborative learning from multi-ordered partitioned neighborhoods and asymmetry-preserving regularizers. Our resulting formalism, Digraph Hyperbolic Networks (D-HYPR), is versatile for multiple tasks including node classification, link presence prediction, and link property prediction. The efficacy of D-HYPR was meticulously examined against 21 previous techniques, using 8 real-world digraph datasets. D-HYPR statistically significantly outperforms the current state of the art. We release our code at https://github. com/hongluzhou/dhypr. Honglu Zhou, Advith Chegu, Samuel S. Sohn, Zuohui Fu, Gerard de Melo, Mubbasir Kapadia |
IJCAI | 6 |
| 2023 | The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational AgentsabstractPrevious studies regarding the perception of emotions for embodied virtual agents have shown the effectiveness of using virtual characters in conveying emotions through interactions with humans. However, creating an autonomous embodied conversational agent with expressive behaviors presents two major challenges. The first challenge is the difficulty of synthesizing the conversational behaviors for each modality that are as expressive as real human behaviors. The second challenge is that the affects are modeled independently, which makes it difficult to generate multimodal responses with consistent emotions across all modalities. In this work, we propose a conceptual framework, ACTOR (Affect-Consistent mulTimodal behaviOR generation), that aims to increase the perception of affects by generating multimodal behaviors conditioned on a consistent driving affect. We have conducted a user study with 199 participants to assess how the average person judges the affects perceived from multimodal behaviors that are consistent and inconsistent with respect to a driving affect. The result shows that among all model conditions, our affect-consistent framework receives the highest Likert scores for the perception of driving affects. Our statistical analysis suggests that making a modality affect-inconsistent significantly decreases the perception of driving affects. We also observe that multimodal behaviors conditioned on consistent affects are more expressive compared to behaviors with inconsistent affects. Therefore, we conclude that multimodal emotion conditioning and affect consistency are vital to enhancing the perception of affects for embodied conversational agents. Che-Jui Chang, Samuel S. Sohn, Rajath Jayashankar, Muhammad Usman 0010, Mubbasir Kapadia |
IUI | 6 |
| 2023 | Collective Intelligence during Emergency Egress: The Mechanisms Underlying Altruistic Information ExchangeabstractUnderstanding the human factors governing effective information exchange is increasingly indispensable for the design of day to day human-computer systems. Moreover, effective information exchange becomes a matter of life or death during emergency egress. The complexity of an unknown environment and the unpredictable locations of hazards often prevent evacuees from identifying safe routes. Successful evacuations from locations impacted by fire or earthquakes may depend on user-generated information to increase the chance of collective survival. The present paper employed multi-user virtual reality experiments and an online survey to investigate the mechanisms underlying social influence and collective intelligence during emergencies. Our results demonstrate that information sharing helps to reduce evacuation time and trajectory length. Participants also shared more when given incentives or when there was a lack of knowledge in the public information pool. This work provides further indications of how collective intelligence can be promoted and deployed during emergencies. Hantao Zhao, Tyler Thrash, Fabian Schläfli, Mubbasir Kapadia, Leonel Aguilar Melgar, Dirk Helbing, Christoph Hölscher |
Int. J. Hum. Comput. Interact. | 4 |
| 2023 | Cognitive Path Planning With Spatial Memory DistortionabstractHuman path-planning operates differently from deterministic AI-based path-planning algorithms due to the decay and distortion in a human's spatial memory and the lack of complete scene knowledge. Here, we present a cognitive model of path-planning that simulates human-like learning of unfamiliar environments, supports systematic degradation in spatial memory, and distorts spatial recall during path-planning. We propose a Dynamic Hierarchical Cognitive Graph (DHCG) representation to encode the environment structure by incorporating two critical spatial memory biases during exploration: categorical adjustment and sequence order effect. We then extend the "Fine-To-Coarse" (FTC), the most prevalent path-planning heuristic, to incorporate spatial uncertainty during recall through the DHCG. We conducted a lab-based Virtual Reality (VR) experiment to validate the proposed cognitive path-planning model and made three observations: (1) a statistically significant impact of sequence order effect on participants' route-choices, (2) approximately three hierarchical levels in the DHCG according to participants' recall data, and (3) similar trajectories and significantly similar wayfinding performances between participants and simulated cognitive agents on identical path-planning tasks. Furthermore, we performed two detailed simulation experiments with different FTC variants on a Manhattan-style grid. Experimental results demonstrate that the proposed cognitive path-planning model successfully produces human-like paths and can capture human wayfinding's complex and dynamic nature, which traditional AI-based path-planning algorithms cannot capture. Rohit Kumar Dubey, Samuel S. Sohn, Tyler Thrash, Christoph Hölscher, André Borrmann, Mubbasir Kapadia |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Heterogeneous Crowd Simulation Using Parametric Reinforcement LearningabstractAgent-based synthetic crowd simulation affords the cost-effective large-scale simulation and animation of interacting digital humans. Model-based approaches have successfully generated a plethora of simulators with a variety of foundations. However, prior approaches have been based on statically defined models predicated on simplifying assumptions, limited video-based datasets, or homogeneous policies. Recent works have applied reinforcement learning to learn policies for navigation. However, these approaches may learn static homogeneous rules, are typically limited in their generalization to trained scenarios, and limited in their usability in synthetic crowd domains. In this article, we present a multi-agent reinforcement learning-based approach that learns a parametric predictive collision avoidance and steering policy. We show that training over a parameter space produces a flexible model across crowd configurations. That is, our goal-conditioned approach learns a parametric policy that affords heterogeneous synthetic crowds. We propose a model-free approach without centralization of internal agent information, control signals, or agent communication. The model is extensively evaluated. The results show policy generalization across unseen scenarios, agent parameters, and out-of-distribution parameterizations. The learned model has comparable computational performance to traditional methods. Qualitatively the model produces both expected (laminar flow, shuffling, bottleneck) and unexpected (side-stepping) emergent qualitative behaviours, and quantitatively the approach is performant across measures of movement quality. Kaidong Hu, M. Brandon Haworth, Glen Berseth, Vladimir Pavlovic 0001, Petros Faloutsos, Mubbasir Kapadia |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Cross-Modal Coherence for Text-to-Image RetrievalabstractCommon image-text joint understanding techniques presume that images and the associated text can universally be characterized by a single implicit model. However, co-occurring images and text can be related in qualitatively different ways, and explicitly modeling it could improve the performance of current joint understanding models. In this paper, we train a Cross-Modal Coherence Model for text-to-image retrieval task. Our analysis shows that models trained with image–text coherence relations can retrieve images originally paired with target text more often than coherence-agnostic models. We also show via human evaluation that images retrieved by the proposed coherence-aware model are preferred over a coherence-agnostic baseline by a huge margin. Our findings provide insights into the ways that different modalities communicate and the role of coherence relations in capturing commonsense inferences in text and imagery. Malihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia, Vladimir Pavlovic 0001, Matthew Stone |
AAAI | 4 |
| 2022 | Impact of Manikin Display on Perception of Spatial PlanningabstractThe visualization of spaces, both virtual and built, has long been an important part of the environment design process. Industry tools to visualize occupancy have grown from simple drop-in stock photos post-design to real-time crowds simulations. However, while treatment of visualization and collaborative design processes has long been discussed in the HCI and Architecture communities, these inclusive design methods are infrequently seen in architecture education (e.g. studio) and practice, nor implemented in licensure requirements – leaving designers to think about the future occupants on their own. While there are strong indicators of the impact visualization modality and rendering style have on perception of scale and space, little has been explored regarding how we represent the human form with respect to these design tools and practices. We present findings from a novel online interactive space planning and estimation study that examines the effects of 3 common building visualization modalities in the design process with 3 human form modalities extracted from the architecture literature. Results indicate the type of visualization changes the number of occupants estimated, and that designers prefer integrated manikins within building models when estimating space usage, although their acceptance was equally divided between 2D and 3D. Our findings lay the foundation for new and focused design tools integrating human form and factors at building scale. Mathew Schwartz, M. Brandon Haworth, Muhammad Usman 0010, Petros Faloutsos, Mubbasir Kapadia |
SAP | 5 |
| 2022 | D-HYPR: Harnessing Neighborhood Modeling and Asymmetry Preservation for Digraph Representation LearningabstractDigraph Representation Learning (DRL) aims to learn representations for directed homogeneous graphs (digraphs). Prior work in DRL is largely constrained (e.g., limited to directed acyclic graphs), or has poor generalizability across tasks (e.g., evaluated solely on one task). Most Graph Neural Networks (GNNs) exhibit poor performance on digraphs due to the neglect of modeling neighborhoods and preserving asymmetry. In this paper, we address these notable challenges by leveraging hyperbolic collaborative learning from multi-ordered and partitioned neighborhoods, and regularizers inspired by socio-psychological factors. Our resulting formalism, Digraph Hyperbolic Networks (D-HYPR) -- albeit conceptually simple -- generalizes to digraphs where cycles and non-transitive relations are common, and is applicable to multiple downstream tasks including node classification, link presence prediction, and link property prediction. In order to assess the effectiveness of D-HYPR, extensive evaluations were performed across 8 real-world digraph datasets involving 21 prior techniques. D-HYPR statistically significantly outperforms the current state of the art. We release our code at https://github.com/hongluzhou/dhypr Honglu Zhou, Advith Chegu, Samuel S. Sohn, Zuohui Fu, Gerard de Melo, Mubbasir Kapadia |
CIKM | 6 |
| 2022 | How Optimal is Too Optimal? Expectations About Performance in the Traveling Salesman Problem
Serena De Stefani, Samuel S. Sohn, Adeeb Kabir, Mubbasir Kapadia, Jacob Feldman |
CogSci | 4 |
| 2022 | A Computational Method for the Classification of Mental Representations of Objects in 3D Space (Short Paper)
Samuel S. Sohn, Panagiotis Mavros, Mubbasir Kapadia, Christoph Hölscher |
COSIT | 3 |
| 2022 | MUSE-VAE: Multi-Scale VAE for Environment-Aware Long Term Trajectory PredictionabstractAccurate long-term trajectory prediction in complex scenes, where multiple agents (e.g., pedestrians or vehicles) interact with each other and the environment while attempting to accomplish diverse and often unknown goals, is a challenging stochastic forecasting problem. In this work, we propose MUSEVAE, a new probabilistic modeling framework based on a cascade of Conditional VAEs, which tackles the long-term, uncertain trajectory prediction task using a coarse-to-fine multi-factor forecasting architecture. In its Macro stage, the model learns a joint pixel-space representation of two key factors, the underlying environment and the agent movements, to predict the long and short term motion goals. Conditioned on them, the Micro stage learns a fine-grained spatio-temporal representation for the prediction of individual agent trajectories. The VAE backbones across the two stages make it possible to naturally account for the joint uncertainty at both levels of granularity. As a result, MUSEVAE offers diverse and simultaneously more accurate predictions compared to the current state-of-the-art. We demonstrate these assertions through a comprehensive set of experiments on nuScenes and SDD benchmarks as well as PFSD, a new synthetic dataset, which challenges the forecasting ability of models on complex agent-environment interaction scenarios. Mihee Lee, Samuel S. Sohn, Seonghyeon Moon, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001 |
CVPR | 5 |
| 2022 | HM: Hybrid Masking for Few-Shot Segmentation
Seonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon, Vladimir Pavlovic 0001, Muhammad Haris Khan, Mubbasir Kapadia |
ECCV (20) | 7 |
| 2022 | COMPOSER: Compositional Reasoning of Group Activity in Videos with Keypoint-Only Modality
Honglu Zhou, Asim Kadav, Aviv Shamsian, Shijie Geng, Farley Lai, Long Zhao 0003, Ting Liu 0005, Mubbasir Kapadia, Hans Peter Graf |
ECCV (35) | 8 |
| 2022 | The IVI Lab entry to the GENEA Challenge 2022 - A Tacotron2 Based Method for Co-Speech Gesture Generation With Locality-Constraint Attention MechanismabstractThis paper describes the IVI Lab entry to the GENEA Challenge 2022. We formulate the gesture generation problem as a sequence-to-sequence conversion task with text, audio, and speaker identity as inputs and the body motion as the output. We use the Tacotron2 architecture as our backbone with the locality-constraint attention mechanism that guides the decoder to learn the dependencies from the neighboring latent features. The collective evaluation released by GENEA Challenge 2022 indicates that our two entries (FSH and USK) for the full body and upper body tracks statistically outperform the audio-driven and text-driven baselines on both two subjective metrics. Remarkably, our full-body entry receives the highest speech appropriateness (60.5% matched) among all submitted entries. We also conduct an objective evaluation to compare our motion acceleration and jerk with two autoregressive baselines. The result indicates that the motion distribution of our generated gestures is much closer to the distribution of natural gestures. Che-Jui Chang, Mubbasir Kapadia |
ICMI | 3 |
| 2022 | Harnessing Fourier Isovists and Geodesic Interaction for Long-Term Crowd Flow PredictionabstractWith the rise in popularity of short-term Human Trajectory Prediction (HTP), Long-Term Crowd Flow Prediction (LTCFP) has been proposed to forecast crowd movement in large and complex environments. However, the input representations, models, and datasets for LTCFP are currently limited. To this end, we propose Fourier Isovists, a novel input representation based on egocentric visibility, which consistently improves all existing models. We also propose GeoInteractNet (GINet), which couples the layers between a multi-scale attention network (M-SCAN) and a convolutional encoder-decoder network (CED). M-SCAN approximates a super-resolution map of where humans are likely to interact on the way to their goals and produces multi-scale attention maps. The CED then uses these maps in either its encoder's inputs or its decoder's attention gates, which allows GINet to produce super-resolution predictions with substantially higher accuracy than existing models even with Fourier Isovists. In order to evaluate the scalability of models to large and complex environments, which the only existing LTCFP dataset is unsuitable for, a new synthetic crowd dataset with both real and synthetic environments has been generated. In its nascent state, LTCFP has much to gain from our key contributions. The Supplementary Materials, dataset, and code are available at sssohn.github.io/GeoInteractNet. Samuel S. Sohn, Seonghyeon Moon, Honglu Zhou, Mihee Lee, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
IJCAI | 7 |
| 2022 | Automatic estimation of parametric saliency maps (PSMs) for autonomous pedestrians
Melissa Kremer, Peter Caruana, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
Comput. Graph. | 4 |
| 2022 | A2X: An end-to-end framework for assessing agent and environment interactions in multimodal human trajectory prediction
Samuel S. Sohn, Mihee Lee, Seonghyeon Moon, Gang Qiao, Muhammad Usman 0010, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
Comput. Graph. | 8 |
| 2022 | Disentangling audio content and emotion with adaptive instance normalization for expressive facial animation synthesisabstractAbstract 3D facial animation synthesis from audio has been a focus in recent years. However, most existing literature works are designed to map audio and visual content, providing limited knowledge regarding the relationship between emotion in audio and expressive facial animation. This work generates audio‐matching facial animations with the specified emotion label. In such a task, we argue that separating the content from audio is indispensable—the proposed model must learn to generate facial content from audio content while expressions from the specified emotion. We achieve it by an adaptive instance normalization module that isolates the content in the audio and combines the emotion embedding from the specified label. The joint content‐emotion embedding is then used to generate 3D facial vertices and texture maps. We compare our method with state‐of‐the‐art baselines, including the facial segmentation‐based and voice conversion‐based disentanglement approaches. We also conduct a user study to evaluate the performance of emotion conditioning. The results indicate that our proposed method outperforms the baselines in animation quality and expression categorization accuracy. Che-Jui Chang, Long Zhao 0003, Mubbasir Kapadia |
Comput. Animat. Virtual Worlds | 4 |
| 2022 | Graph-based generative representation learning of semantically and behaviorally augmented floorplans
Vahid Azizi 0003, Muhammad Usman 0010, Honglu Zhou, Petros Faloutsos, Mubbasir Kapadia |
Vis. Comput. | 5 |
| 2021 | The Interplay Between Local and Global Strategies in Navigational Decisions
Serena De Stefani, Samuel S. Sohn, Adeeb Kabir, Mubbasir Kapadia, Jacob Feldman |
CogSci | 4 |
| 2021 | AESOP: Abstract Encoding of Stories, Objects, and Pictures
Hareesh Ravi, Kushal Kafle, Scott Cohen, Jonathan Brandt, Mubbasir Kapadia |
ICCV | 5 |
| 2021 | Hopper: Multi-hop Transformer for Spatiotemporal Reasoning
Honglu Zhou, Asim Kadav, Farley Lai, Alexandru Niculescu-Mizil, Martin Renqiang Min, Mubbasir Kapadia, Hans Peter Graf |
ICLR | 6 |
| 2021 | SNAP: Successor Entropy based Incremental Subgoal Discovery for Adaptive NavigationabstractReinforcement learning (RL) has demonstrated great success in solving navigation tasks but often fails when learning complex environmental structures. One open challenge is to incorporate low-level generalizable skills with human-like adaptive path-planning in an RL framework. Motivated by neural findings in animal navigation, we propose a Successor eNtropy-based Adaptive Path-planning (SNAP) that combines a low-level goal-conditioned policy with the flexibility of a classical high-level planner. SNAP decomposes distant goal-reaching tasks into multiple nearby goal-reaching sub-tasks using a topological graph. To construct this graph, we propose an incremental subgoal discovery method that leverages the highest-entropy states in the learned Successor Representation. The Successor Representation encodes the likelihood of being in a future state given the current state and capture the relational structure of states based on a policy. Our main contributions lie in discovering subgoal states that efficiently abstract the state-space and proposing a low-level goal-conditioned controller for local navigation. Since the basic low-level skill is learned independent of state representation, our model easily generalizes to novel environments without intensive relearning. We provide empirical evidence that the proposed method enables agents to perform long-horizon sparse reward tasks quickly, take detours during barrier tasks, and exploit shortcuts that did not exist during training. Our experiments further show that the proposed method outperforms the existing goal-conditioned RL algorithms in successfully reaching distant-goal tasks and policy learning. To evaluate human-like adaptive path-planning, we also compare our optimal agent with human data and found that, on average, the agent was able to find a shorter path than the human participants. Rohit Kumar Dubey, Samuel S. Sohn, Jimmy Abualdenien, Tyler Thrash, Christoph Hölscher, André Borrmann, Mubbasir Kapadia |
MIG | 7 |
| 2021 | PSM: Parametric Saliency Maps for Autonomous PedestriansabstractModeling visual attention is an important aspect of simulating realistic virtual humans. This work proposes a parametric model and method for generating real-time saliency maps from the perspective of virtual agents which approximate those of vision-based saliency approaches. The model aggregates a saliency score from user-defined parameters for objects and characters in an agent’s view and uses that to output a 2D saliency map which can be modulated by an attention field to incorporate 3D information as well as a character’s state of attentiveness. The aggregate and parameterized structure of the method allows the user to model a range of diverse agents. The user may also expand the model with additional layers and parameters. The proposed method can be combined with normative and pathological models of the human visual field and gaze controllers, such as the recently proposed model of egocentric distractions for casual pedestrians that we use in our results. Melissa Kremer, Peter Caruana, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 4 |
| 2021 | A2X: An Agent and Environment Interaction Benchmark for Multimodal Human Trajectory PredictionabstractIn recent years, human trajectory prediction (HTP) has garnered attention in computer vision literature. Although this task has much in common with the longstanding task of crowd simulation, there is little from crowd simulation that has been borrowed, especially in terms of evaluation protocols. The key difference between the two tasks is that HTP is concerned with forecasting multiple steps at a time and capturing the multimodality of real human trajectories. A majority of HTP models are trained on the same few datasets, which feature small, transient interactions between real people and little to no interaction between people and the environment. Unsurprisingly, when tested on crowd egress scenarios, these models produce erroneous trajectories that accelerate too quickly and collide too frequently, but the metrics used in HTP literature cannot convey these particular issues. To address these challenges, we propose (1) the A2X dataset, which has simulated crowd egress and complex navigation scenarios that compensate for the lack of agent-to-environment interaction in existing real datasets, and (2) evaluation metrics that convey model performance with more reliability and nuance. A subset of these metrics are novel multiverse metrics, which are better-suited for multimodal models than existing metrics. The dataset is available at: https://mubbasir.github.io/HTP-benchmark/. Samuel S. Sohn, Mihee Lee, Seonghyeon Moon, Gang Qiao, Muhammad Usman 0010, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
MIG | 8 |
| 2021 | Simulation-as-a-Service: Analyzing Crowd Movements in Virtual EnvironmentsabstractAbstract At present, environment designers mostly use their intuition and experience to predictively account for how environments might support dynamic activity. The majority of Computer‐Aided Design tools only provide a static representation of space which potentially ignores the impact that an environment layout produces on its occupants and their movements. To address this, computational techniques such as crowd simulation have been developed. With few exceptions, crowd simulation frameworks are often decoupled from environment modeling tools. They usually require specific hardware/software infrastructures and expertise to be used, hindering the designers' abilities to seamlessly simulate, analyze, and incorporate movement‐centric dynamics into their design workflows. To bridge this disconnect, we devise a cross‐browser service‐based simulation analytics platform to analyze environment layouts with respect to occupancy and activity. Our platform allows users to access simulation services by uploading three‐dimensional environment models in numerous common formats, devise targeted simulation scenarios, run simulations, and instantly generate crowd‐based analytics for their designs. We conducted a case study to showcase cross‐domain applicability of our service‐based platform, and a user study to evaluate the usability of this approach. Muhammad Usman 0010, M. Brandon Haworth, Petros Faloutsos, Mubbasir Kapadia |
Comput. Animat. Virtual Worlds | 4 |
| 2021 | Interactive Architectural Design with Diverse Solution ExplorationabstractIn architectural design, architects explore a vast amount of design options to maximize various performance criteria, while adhering to specific constraints. In an effort to assist architects in such a complex endeavour, we propose IDOME, an interactive system for computer-aided design optimization. Our approach balances automation and control by efficiently exploring, analyzing, and filtering space layouts to inform architects' decision-making better. At each design iteration, IDOME provides a set of alternative building layouts which satisfy user-defined constraints and optimality criteria concerning a user-defined space parametrization. When the user selects a design generated by IDOME, the system performs a similar optimization process with the same (or different) parameters and objectives. A user may iterate this exploration process as many times as needed. In this work, we focus on optimizing built environments using architectural metrics by improving the degree of visibility, accessibility, and information gaining for navigating a proposed space. This approach, however, can be extended to support other kinds of analysis as well. We demonstrate the capabilities of IDOME through a series of examples, performance analysis, user studies, and a usability test. The results indicate that IDOME successfully optimizes the proposed designs concerning the chosen metrics and offers a satisfactory experience for users with minimal training. Glen Berseth, M. Brandon Haworth, Muhammad Usman 0010, Davide Schaumann, Mahyar Khayatkhoei, Mubbasir Kapadia, Petros Faloutsos |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Modelling distracted agents in crowd simulations
Melissa Kremer, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
Vis. Comput. | 3 |
| 2020 | Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior KnowledgeabstractCross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student). Existing approaches require paired training examples exist in both modalities. However, accessing the data from superior modalities may not always be feasible. For example, in the case of 3D hand pose estimation, depth maps, point clouds, or stereo images usually capture better hand structures than RGB images, but most of them are expensive to be collected. In this paper, we propose a novel scheme to train the Student in a Target dataset where the Teacher is unavailable. Our key idea is to generalize the distilled cross-modal knowledge learned from a Source dataset, which contains paired examples from both modalities, to the Target dataset by modeling knowledge as priors on parameters of the Student. We name our method "Cross-Modal Knowledge Generalization" and demonstrate that our scheme results in competitive performance for 3D hand pose estimation on standard benchmark datasets. Long Zhao 0003, Xi Peng 0005, Yuxiao Chen 0002, Mubbasir Kapadia, Dimitris N. Metaxas |
CVPR | 4 |
| 2020 | Laying the Foundations of Deep Long-Term Crowd Flow Prediction
Samuel S. Sohn, Honglu Zhou, Seonghyeon Moon, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
ECCV (29) | 6 |
| 2020 | HID: Hierarchical Multiscale Representation Learning for Information DiffusionabstractMultiscale modeling has yielded immense success on various machine learning tasks. However, it has not been properly explored for the prominent task of information diffusion, which aims to understand how information propagates along users in online social networks. For a specific user, whether and when to adopt a piece of information propagated from another user is affected by complex interactions, and thus, is very challenging to model. Current state-of-the-art techniques invoke deep neural models with vector representations of users. In this paper, we present a Hierarchical Information Diffusion (HID) framework by integrating user representation learning and multiscale modeling. The proposed framework can be layered on top of all information diffusion techniques that leverage user representations, so as to boost the predictive power and learning efficiency of the original technique. Extensive experiments on three real-world datasets showcase the superiority of our method. Honglu Zhou, Zuohui Fu, Gerard de Melo, Yongfeng Zhang 0003, Mubbasir Kapadia |
IJCAI | 6 |
| 2020 | Deep Integration of Physical Humanoid Control and Crowd NavigationabstractMany multi-agent navigation approaches make use of simplified representations such as a disk. These simplifications allow for fast simulation of thousands of agents but limit the simulation accuracy and fidelity. In this paper, we propose a fully integrated physical character control and multi-agent navigation method. In place of sample complex online planning methods, we extend the use of recent deep reinforcement learning techniques. This extension improves on multi-agent navigation models and simulated humanoids by combining Multi-Agent and Hierarchical Reinforcement Learning. We train a single short term goal-conditioned low-level policy to provide directed walking behaviour. This task-agnostic controller can be shared by higher-level policies that perform longer-term planning. The proposed approach produces reciprocal collision avoidance, robust navigation, and emergent crowd behaviours. Furthermore, it offers several key affordances not previously possible in multi-agent navigation including tunable character morphology and physically accurate interactions with agents and the environment. Our results show that the proposed method outperforms prior methods across environments and tasks, as well as, performing well in terms of zero-shot generalization over different numbers of agents and computation time. M. Brandon Haworth, Glen Berseth, Seonghyeon Moon, Petros Faloutsos, Mubbasir Kapadia |
MIG | 5 |
| 2020 | Watch Out! Modelling Pedestrians with Egocentric DistractionsabstractThe use of mobile devices is one of the most commonly observed family of distracted behaviours exhibited by pedestrians in urban environments. We develop an event-driven behaviour tree model for distracted pedestrians that includes initiating mobile device use as well as terminating or pausing mobile device use based on internal or external cues to refocus attention. We present a simple, probabilistic attention model for such pedestrians. The proposed model is not meant to be complete. It primarily focuses on computing the probability that a distracted agent looks up, based on the agent’s individual characteristics and the elements in their environment. We condition the potentially attention grabbing elements in the environment on distraction-specific egocentric fields for visual attention. We also propose an oriented ellipse model for capturing the affects of cognitively fuzzy goals during distracted navigation. Our model is simple and intuitively parameterized, and thus can be easily edited and extended. Melissa Kremer, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 3 |
| 2020 | A Social Distancing Index: Evaluating Navigational Policies on Human Proximity using Crowd SimulationsabstractThe importance of social distancing for public health is well established. However, the policies and regulations regarding occupancy rates have not been designed with this in mind. While there are analytical tools and related measures that are used in practice to evaluate how the design of a built environment serves the needs of its intended occupants, these metrics cannot directly apply to the problem of preventing the spread of infectious diseases such as COVID-19. By using a crowd-based simulator using three levels of behavior and agent control in a given environment, a novel evaluation metric for a space layout can be calculated to reflect the proclivity of maintaining a safe distance throughout the shopping experience. We refer to this metric as the Social Distancing Index (SDI), accounting for the occupancy throughput and number of distance-based violations found. Through a case study of a realistic retail store, we demonstrate the proposed platforms performance and output on multiple scenarios by changing agent-behavior, occupancy rate, and navigational guidelines. Muhammad Usman 0010, Tien-Chi Lee, Ryhan Moghe, Petros Faloutsos, Mubbasir Kapadia |
MIG | 6 |
| 2020 | AUTOSIGN: A multi-criteria optimization approach to computer aided design of signage layouts in complex buildings
Rohit Kumar Dubey, Wei Ping Khoo, Michal Gath-Morad, Christoph Hölscher, Mubbasir Kapadia |
Comput. Graph. | 5 |
| 2020 | Predicting Crowd Egress and Environment Relationships to Support Building Design Optimization
Kaidong Hu, Sejong Yoon, Vladimir Pavlovic 0001, Petros Faloutsos, Mubbasir Kapadia |
Comput. Graph. | 5 |
| 2020 | Towards Image-to-Video Translation: A Structure-Aware Approach via Multi-stage Generative Adversarial Networks
Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas |
Int. J. Comput. Vis. | 4 |
| 2020 | Generation of crowd arrival and destination locations/times in complex transit facilitiesabstractAbstract In order to simulate virtual agents in the replica of a real facility across a long time span, a crowd simulation engine needs a list of agent arrival and destination locations and times that reflect those seen in the actual facility. Working together with a major metropolitan transportation authority, we propose a specification that can be used to procedurally generate this information. This specification is both uniquely compact and expressive—compact enough to mirror the mental model of building managers and expressive enough to handle the wide variety of crowds seen in real urban environments. We also propose a procedural algorithm for generating tens of thousands of high-level agent paths from this specification. This algorithm allows our specification to be used with traditional crowd simulation obstacle avoidance algorithms while still maintaining the realism required for the complex, real-world simulations of a transit facility. Our evaluation with industry professionals shows that our approach is intuitive and provides controls at the right level of detail to be used in large facilities (200,000+ people/day). Brian Ricks, Andrew Dobson, Athanasios Krontiris, Kostas E. Bekris, Mubbasir Kapadia, Fred S. Roberts |
Vis. Comput. | 5 |
| 2019 | Semantic Graph Convolutional Networks for 3D Human Pose RegressionabstractIn this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convolution filters and shared transformation matrix for each node. To address these limitations, we propose Semantic Graph Convolutional Networks (SemGCN), a novel neural network architecture that operates on regression tasks with graph-structured data. SemGCN learns to capture semantic information such as local and global node relationships, which is not explicitly represented in the graph. These semantic relationships can be learned through end-to-end training from the ground truth without additional supervision or hand-crafted rules. We further investigate applying SemGCN to 3D human pose regression. Our formulation is intuitive and sufficient since both 2D and 3D human poses can be represented as a structured graph encoding the relationships between joints in the skeleton of a human body. We carry out comprehensive studies to validate our method. The results prove that SemGCN outperforms state of the art while using 90% fewer parameters. Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas |
CVPR | 4 |
| 2019 | Fusion-Based Wayfinding Prediction Model for Multiple Information Sources
Rohit Kumar Dubey, Samuel S. Sohn, Christoph Hölscher, Mubbasir Kapadia |
FUSION | 4 |
| 2019 | JUNGLE: An Interactive Visual Platform for Collaborative Creation and Consumption of Nonlinear Transmedia Stories
Mubbasir Kapadia, Carlos Muñiz 0001, Samuel S. Sohn, Sasha Schriber, Kenny Mitchell, Markus Gross 0001 |
ICIDS | 1 |
| 2019 | StoryPrint: an interactive visualization of storiesabstractIn this paper, we propose StoryPrint, an interactive visualization of creative storytelling that facilitates individual and comparative structural analyses. This visualization method is intended for script-based media, which has suitable metadata. The pre-visualization process involves parsing the script into different metadata categories and analyzing the sentiment on a character and scene basis. For each scene, the setting, character presence, character prominence, and character emotion of a film are represented as a StoryPrint. The visualization is presented as a radial diagram of concentric rings wrapped around a circular time axis. A user then has the ability to toggle a difference overlay to assist in the cross-comparison of two different scene inputs. Katie Watson, Samuel S. Sohn, Sasha Schriber, Markus Gross 0001, Carlos Muñiz 0001, Mubbasir Kapadia |
IUI | 6 |
| 2019 | An Interdependent Model of Personality, Motivation, Emotion, and Mood for Intelligent Virtual AgentsabstractBuilding intelligent agents that can believably interact with humans is a difficult yet important task in a host of applications, including therapy, education, and entertainment. We submit that in order to enhance believability, the agent's affective state should be accurately modeled and should realistically influence the agent's behavior. We propose a computational model of affect which incorporates an empirically-based interplay between its various affective components - personality, motivation, emotion, and mood. Further, our model captures a number of salient mechanisms that are observable in humans and that influence the agent's behavior. We are therefore hopeful that our model will facilitate more engaging and meaningful human-agent interactions. We evaluate our model and illustrate its efficacy, as well as the importance of the different components in the model and their interplay. Maayan Shvo, Jakob Buhmann, Mubbasir Kapadia |
IVA | 3 |
| 2019 | Towards a Conversational Interface for Authoring Intelligent Virtual CharactersabstractThe collaboration between creatives and domain architects is crucial for bringing virtual characters to life. Domain architects are technical experts who are tasked with formally designing intelligent virtual characters' domain knowledge, which is a symbolic representation of knowledge that the character uses to reason over its interactions with other agents. In the context of this work, domain knowledge encompasses the mental modeling of the character. Although the creation of interactive narratives requires substantial engineering expertise, it is also necessary to pick the brains of writers, artists, and animators alike to give the characters a boost of peculiarities. This intrinsically collaborative and interdisciplinary process brings about the challenge of bridging different mindsets and workflows in an efficient and effective way. Samuel S. Sohn, Mubbasir Kapadia |
IVA | 3 |
| 2019 | Joint Exploration and Analysis of High-Dimensional Design-Occupancy TemplatesabstractCrowd simulations provide a practical approach to evaluate building design alternatives with respect to human-centric criteria, such as evacuation times and flow in case of emergency scenarios. Coupled with Building Information Modeling (BIM) tools, they support architects’ iterative exploration of design alternatives. However, methods based on manually configuring a design and a corresponding simulation are not practical for exploring the potentially very large number of design solutions that satisfy human-centric design goals and requirements. Often, for practical reasons, designers may consider standard crowd configurations which do not capture the behavior of diverse occupants that may exhibit different locomotion abilities, movement patterns, and social behaviors. We posit that a joint exploration of high-dimensional building design and occupancy features is necessary to more accurately capture the mutual relations between buildings and the behavior of their occupants. To test this hypothesis, we conducted a series of experiments to automatically explore joint high dimensional design–occupancy patterns using an unsupervised pattern recognition technique (i.e. K-MEANS). We demonstrate that joint design–occupancy explorations provide more accurate results compared with sequential exploration processes that consider default design or crowd features, despite the longer computational times to simulate a large number of solutions. The findings of this case study have practical applications to the design of next-generation design exploration tools that support human-centric analyses in architectural design. Muhammad Usman 0010, Davide Schaumann, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 4 |
| 2019 | Identifying Indoor Navigation Landmarks Using a Hierarchical Multi-Criteria Decision FrameworkabstractLandmarks play a vital role in human wayfinding by providing the structure for mental spatial representations and indicating locations with which to orient. Less research effort has been allocated towards automated landmark identification in indoor environments despite a growing interest in indoor navigation in the scientific community. In this paper, we propose a computational framework to identify indoor landmarks that is based on a hierarchical multi-criteria decision model and grounded in theories of spatial cognition and human information processing. Our model of landmark salience is represented as a hierarchical integration process of low-level features derived from a three-part, higher-level, salience vector (i.e., cognitive, spatial, and subjective salience). We use a fuzzy hierarchical composite-weighted (objective and subjective) Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) to derive the rankings for identified objects at decision points (i.e., intersections). The top N objects are then selected and compared to a list of landmarks derived from an eye-tracking based virtual reality (VR) experiment. A substantial overlap of 79% was observed between these two lists. The proposed framework is capable of reliably and accurately detecting indoor landmarks, which can be employed in the development of landmark-based robot/autonomous agent motion and indoor guidance systems. Rohit Kumar Dubey, Samuel S. Sohn, Tyler Thrash, Christoph Hölscher, Mubbasir Kapadia |
MIG | 5 |
| 2019 | Scenario Generalization of Data-driven Imitation Models in Crowd SimulationabstractCrowd simulation, the study of the movement of multiple agents in complex environments, presents a unique application domain for machine learning. One challenge in crowd simulation is to imitate the movement of expert agents in highly dense crowds. An imitation model could substitute an expert agent if the model behaves as good as the expert. This will bring many exciting applications. However, we believe no prior studies have considered the critical question of how training data and training methods affect imitators when these models are applied to novel scenarios. In this work, a general imitation model is represented by applying either the Behavior Cloning (BC) training method or a more sophisticated Generative Adversarial Imitation Learning (GAIL) method, on three typical types of data domains: standard benchmarks for evaluating crowd models, random sampling of state-action pairs, and egocentric scenarios that capture local interactions. Simulated results suggest that (i) simpler training methods are overall better than more complex training methods, (ii) training samples with diverse agent-agent and agent-obstacle interactions are beneficial for reducing collisions when the trained models are applied to new scenarios. We additionally evaluated our models in their ability to imitate real world crowd trajectories observed from surveillance videos. Our findings indicate that models trained on representative scenarios generalize to new, unseen situations observed in real human crowds. Gang Qiao, Honglu Zhou, Mubbasir Kapadia, Sejong Yoon, Vladimir Pavlovic 0001 |
MIG | 3 |
| 2019 | Cartoonish sketch-based face editing in videos using identity deformation transfer
Long Zhao 0003, Fangda Han, Xi Peng 0005, Mubbasir Kapadia, Vladimir Pavlovic 0001, Dimitris N. Metaxas |
Comput. Graph. | 5 |
| 2019 | Coupling agent motivations and spatial behaviors for authoring multiagent narrativesabstractAbstract Authoring behavior narratives for heterogeneous multiagent virtual humans engaged in collaborative, localized, and task‐based behaviors can be challenging. Traditional behavior authoring frameworks are either space‐centric, where occupancy parameters are specified; behavior‐centric, where multiagent behaviors are defined; or agent‐centric, where desires and intentions drive agents' behavior. In this paper, we propose to integrate these approaches into a unique framework to author behavior narratives that progressively satisfy time‐varying building‐level occupancy specifications, room‐level behavior distributions, and agent‐level motivations using a prioritized resource allocation system. This approach can generate progressively more complex and plausible narratives that satisfy spatial, behavioral, and social constraints. Possible applications of this system involve computer gaming and decision‐making in engineering and architectural design. Davide Schaumann, M. Brandon Haworth, Petros Faloutsos, Mubbasir Kapadia |
Comput. Animat. Virtual Worlds | 5 |
| 2018 | The Role of Data-Driven Priors in Multi-Agent Crowd Trajectory EstimationabstractResource constraints frequently complicate multi-agent planning problems. Existing algorithms for resource-constrained, multi-agent planning problems rely on the assumption that the constraints are deterministic. However, frequently resource constraints are themselves subject to uncertainty from external influences. Uncertainty about constraints is especially challenging when agents must execute in an environment where communication is unreliable, making on-line coordination difficult. In those cases, it is a significant challenge to find coordinated allocations at plan time depending on availability at run time. To address these limitations, we propose to extend algorithms for constrained multi-agent planning problems to handle stochastic resource constraints. We show how to factorize resource limit uncertainty and use this to develop novel algorithms to plan policies for stochastic constraints. We evaluate the algorithms on a search-and-rescue problem and on a power-constrained planning domain where the resource constraints are decided by nature. We show that plans taking into account all potential realizations of the constraint obtain significantly better utility than planning for the expectation, while causing fewer constraint violations. Gang Qiao, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001 |
AAAI | 3 |
| 2018 | Computer-Assisted Authoring for Natural Language Story ScriptsabstractIn order to assist scriptwriters during the process of story-writing, we have developed a system that can extract information from natural language stories, and allow for story-centric as well as character-centric reasoning. These inferencing capabilities are exposed to the user through intuitive querying systems, allowing the scriptwriter to ask the system questions about story and character information. We introduce knowledge bytes as atoms of information and demonstrate that the system can parse text into a stream of knowledge bytes and use these mentioned reasoning capabilities through logical reasoning. Rushit Sanghrajka, Wojciech Witon, Sasha Schriber, Markus Gross 0001, Mubbasir Kapadia |
AAAI | 5 |
| 2018 | Efficiency in Solving the Traveling Salesman Problem as Predictor of Perceived Humanness
Serena De Stefani, Samuel S. Sohn, Jacob Feldman, Mubbasir Kapadia, Peter C. Pantelis |
CogSci | 4 |
| 2018 | Show Me a Story: Towards Coherent Neural Story IllustrationabstractWe propose an end-to-end network for visual illustration of a sequence of sentences forming a story. At the core of our model is the ability to model the inter-related nature of the sentences within a story, as well as the ability to learn coherence to support reference resolution. The framework takes the form of an encoder-decoder architecture, where sentences are encoded using a hierarchical two-level sentence-story GRU, combined with an encoding of coherence, and sequentially decoded using a predicted feature representation into a consistent illustrative image sequence. We optimize all parameters of our network in an end-to-end fashion with respect to order embedding loss, encoding entailment between images and sentences. Experiments on the VIST storytelling dataset [9] highlight the importance of our algorithmic choices and efficacy of our overall model. Hareesh Ravi, Lezi Wang, Carlos Muñiz 0001, Leonid Sigal, Dimitris N. Metaxas, Mubbasir Kapadia |
CVPR | 6 |
| 2018 | Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas |
ECCV (15) | 4 |
| 2018 | CARDINAL: Computer Assisted Authoring of Movie ScriptsabstractWe present Cardinal, a tool for computer-assisted authoring of movie scripts. Cardinal provides a means of viewing a script through a variety of perspectives, for interpretation as well as editing. This is made possible by virtue of intelligent automated analysis of natural language scripts and generating different intermediate representations. Cardinal generates 2-D and 3-D visualizations of the scripted narrative and also presents interactions in a timeline-based view. The visualizations empower the scriptwriter to understand their story from a spatial perspective, and the timeline view provides an overview of the interactions in the story. The user study reveals that users of the system demonstrated confidence and comfort using the system. Marcel Marti, Jodok Vieli, Wojciech Witon, Rushit Sanghrajka, Daniel Inversini, Diana Wotruba, Isabel Simo, Sasha Schriber, Mubbasir Kapadia, Markus Gross 0001 |
IUI | 9 |
| 2018 | PICA: Proactive Intelligent Conversational Agent for Interactive NarrativesabstractA narrative relies on the imperfect knowledge of the user to create interactions between the characters that are ultimately used as a plot device to drive the narrative. This motivates our exploration of ways to encode this information, provides means for a user to both query and influence the knowledge, and guides the user based on a model of their experience. We developed PICA: a proactive intelligent conversational agent for interactive narratives that can guide users through such experiences. The underlying knowledge base is designed using a sub-symbolic architecture, which encodes belief models for multiple users and autonomous agents in addition to the actual story knowledge. We also developed a discourse module using Behavior Trees to intuitively design the proactive and reactive capabilities of PICA. We compare our approach to neural networks and symbolic knowledge bases and demonstrate its functionality. Jessica Falk, Steven Poulakos, Mubbasir Kapadia, Robert W. Sumner |
IVA | 3 |
| 2018 | Interactive spatial analytics for human-aware building designabstractWe present a computational spatial analytics tool for designing environments that better support human-related factors. Our system performs both static and dynamic analyses: the first relates to the building geometry and organization, while the second additionally considers the crowd movement in the space. The results are presented to the designers in the form of numerical values, traces and heat maps displayed on top of the floor plan. We demonstrate our approach with a user study whereby novice architects have tested the proposed approach to iteratively improve a building accessibility in real-time with respect to a selected number of static and dynamic metrics. The results indicate that the users were able to successfully improve their design solutions and thus generate more human-aware environments. The usability and effectiveness of the tool where also measured, yielding positive scores. The modular and flexible nature of the tool enables further extension to incorporate additional static and dynamic spatial metrics. Muhammad Usman 0010, Davide Schaumann, M. Brandon Haworth, Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 5 |
| 2018 | Dynamic cognitive maps for agent landmark navigation in unseen environmentsabstractThe development of autonomous agents for wayfinding tasks has long maintained the usage of naive, omniscient models for navigation. The simplicity of these models improves the scalability of crowd simulations, but limits the utility of such simulations to the visualization of general behaviors. This restricted scope does not allow for the observation of more nuanced, individualized behaviors. In this paper, we demonstrate a novel framework for agent simulations that does not rely on omniscience. Instead, each agent is equipped with a memory architecture that enables wayfinding by maintaining a cognitive map of the space explored by the agent. Based on findings from simulation studies, cognitive science, and psychology, we describe a wayfinding procedure that simulates human behavior and human cognitive processes, incorporating landmark navigation, path integration, and memory. This cognitive approach makes observations of agent behavior more comparable to those of human behavior. Samuel S. Sohn, Serena De Stefani, Mubbasir Kapadia |
MIG | 3 |
| 2017 | Perceptual evaluation of space in virtual environmentsabstractFloor plan designs and their spatial analysis are typically constrained to blueprints and 2D projections of 3D models. Computing appropriate spatial measures from such representations provides a standard way of quantifying important aspects of the design. We wish to investigate whether a person's perceptual exploration of a space would agree with such spatial measures, that is, whether a person can roughly infer such measures by exploring a space. We perform two studies, one involving novices and the other experts. First, we conduct a perceptual study to discover whether a novice user's perception of spatial measures depends on the mode used to explore the space. Our analysis considers three spatial measures, grounded in Space-Syntax, that characterize key aspects of a design such as visibility, accessibility, and organization. We compare three modes of exploration: 2D blueprints, first-person view in a 3D simulation, and a 3D virtual reality simulation with teleportation. A correlation analysis between the users' perceptual ratings and the spatial measures, indicates that virtual reality is the most effective of the three methods, while 2D blueprints and 3D first-person exploration often fail entirely to convey the spatial measures. In the second study, experts are asked to evaluate and rank the design blueprints for each measure. The expert observations are in strong agreement with the spatial measures for accessibility and organization, but not for visibility in some cases. This indicates that even experts have difficulty understanding spatial aspects of an architecture design from 2D blueprints alone. Muhammad Usman 0010, M. Brandon Haworth, Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 4 |
| 2017 | Crowd sourced co-design of floor plans using simulation guided gamesabstractCrowd-aware environment design is a complex combinatorial decision process, where small changes in a design may affect crowd flow patterns in unexpected and potentially unintuitive ways. Existing solutions rely on expert intuition, best practices, or automation. To address the dimensionality and complexity of the design process, we propose leveraging automation and human creativity at a large scale akin to crowd sourcing, within a gamified collaborative design framework. Using our system, "players" (novice users or experts) can rapidly iterate on their designs while soliciting feedback from computer simulations of crowd movement and the designs of other players. Our approach affords a new way of thinking of the solution space in that it inherently supports competitive collaboration, co-design, and crowd sourced solutions. We evaluate our framework through a preliminary user study. Nilay Chakraborty, Glen Berseth, M. Brandon Haworth, Petros Faloutsos, Muhammad Usman 0010, Mubbasir Kapadia |
MIG | 6 |
| 2017 | Characterizing the relationship between environment layout and crowd movement using machine learningabstractCrowd simulations facilitate the study of how an environment layout impacts the movement and behavior of its inhabitants. However, simulations are computationally expensive, which make them infeasible when used as part of interactive systems (e.g., Computer-Assisted Design software). Machine learning models, such as neural networks (NN), can learn observed behaviors from examples, and can potentially offer a rational prediction of a crowd's behavior efficiently. To this end, we propose a method to predict the aggregate characteristics of crowd dynamics using regression neural networks (NN). We parametrize the environment, the crowd distribution and the steering method to serve as inputs to the NN models, while a number of common performance measures serve as the output. Our preliminary experiments show that our approach can help users evaluate a large number of environments efficiently. Weining Liu, Vladimir Pavlovic 0001, Kaidong Hu, Petros Faloutsos, Sejong Yoon, Mubbasir Kapadia |
MIG | 6 |
| 2017 | On density-flow relationships during crowd evacuationabstractAbstract Traffic and pedestrian dynamics communities often use a standard qualitative classification, namely, level of service (LoS), to describe the relationship between the crowd flow and crowd density in an environment. However, this classification has not yet been rigorously studied in the application of synthetic crowds, which are derived using a variety of approaches and may model certain behaviors better than others. Although synthetic crowds can be simulated to extrapolate crowd flow for rigorous quantitative analysis, these may be at odds with the qualitative LoS. In order to successfully use computer‐assisted design, it is important to have sound quantitative metrics as the basis for analysis and optimization. In this paper, we present a systematic empirical analysis of LoS for synthetic crowds. Using established crowd simulation techniques, we quantify the relation between crowd density and crowd flow for evacuation scenarios across different simulators to explore conformity to qualitative LoS classifications. Following this study, we perform environment optimization experiments under various LoS conditions. Finally, we test the generality of optimizing under these LoS conditions. Our results motivate the need for further study, using real and synthetic crowd datasets across representative environment benchmarks. M. Brandon Haworth, Muhammad Usman 0010, Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 4 |
| 2017 | CODE: Crowd-optimized design of environmentsabstractAbstract We present crowd‐optimized design of environments (CODE): a “crowd‐aware” computational tool for designing environments (e.g., building floor plans). Our system analyses the impact of newly added environment elements (e.g., pillars or doorways) on the resulting crowd flow, using current‐generation crowd simulators. The results of the simulation are used to provide feedback to the designer in terms of aggregate statistics and heat maps. Additionally, our system is able to “automatically” optimize the placement of environment elements to maximize crowd flow in egress scenarios, while satisfying constraints that are imposed by the designer. Using CODE, architects and environment designers can iteratively refine upon their original design to quickly accommodate the dynamic properties of crowd simulations in an interactive fashion. CODE is modular and flexible so that designers may build environments, select from different crowd simulators, and specify varying crowd configurations. M. Brandon Haworth, Muhammad Usman 0010, Glen Berseth, Mahyar Khayatkhoei, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 5 |
| 2017 | PERFORM: Perceptual Approach for Adding OCEAN Personality to Human Motion Using Laban Movement AnalysisabstractA major goal of research on virtual humans is the animation of expressive characters that display distinct psychological attributes. Body motion is an effective way of portraying different personalities and differentiating characters. The purpose and contribution of this work is to describe a formal, broadly applicable, procedural, and empirically grounded association between personality and body motion and apply this association to modify a given virtual human body animation that can be represented by these formal concepts. Because the body movement of virtual characters may involve different choices of parameter sets depending on the context, situation, or application, formulating a link from personality to body motion requires an intermediate step to assist generalization. For this intermediate step, we refer to Laban Movement Analysis, which is a movement analysis technique for systematically describing and evaluating human motion. We have developed an expressive human motion generation system with the help of movement experts and conducted a user study to explore how the psychologically validated OCEAN personality factors were perceived in motions with various Laban parameters. We have then applied our findings to procedurally animate expressive characters with personality, and validated the generalizability of our approach across different models and animations via another perception study. Funda Durupinar, Mubbasir Kapadia, Susan Deutsch, Michael Neff, Norman I. Badler |
ACM Trans. Graph. | 2 |
| 2016 | ACUMEN: Activity-Centric Crowd Authoring Using Influence MapsabstractHeterogeneity in virtual crowds is crucial for many applications, including visual effects, games, and security simulations. Nevertheless, tweaking the behavior parameters of a character to achieve crowd heterogeneity is frequently hard. In particular, it is typically unclear how tuning some non-intuitive parameters at the agent level will eventually affect both the microscopic or macroscopic scale of the crowd. This paper proposes an activity-centric framework for authoring functional, heterogeneous virtual crowds in semantically meaningful environments. The specification of locations as environmental attractors and agent desires are used to compute "influence maps", which allow the emergence of heterogeneous behaviors in a large virtual crowd in a complex scene. The same framework can also facilitate the authoring of complex group behaviors, such as following behaviors or families, by treating moving agents as attractors. Accompanying results demonstrate the framework's potential by authoring crowds in different environments. The experiments highlight the ability to easily orchestrate purposeful, heterogeneous crowd activities both at a macroscopic and microscopic level with minimal parameter tuning. Athanasios Krontiris, Kostas E. Bekris, Mubbasir Kapadia |
CASA | 3 |
| 2016 | Evaluating Accessible Graphical Interfaces for Building Story Worlds
Steven Poulakos, Mubbasir Kapadia, Guido M. Maiga, Fabio Zünd, Markus Gross 0001, Robert W. Sumner |
ICIDS | 2 |
| 2016 | Confluence: visualizing social physics for interactive narrativeabstractWe introduce Confluence, a web-based social physics game framework and development tools. It seeks to combine the powers of graphical game engines with physics and animation, and social simulation of social physics engines, with convenient tools like social network visualization, strategy analysis and in-game social rule authoring. We evaluate it in developing a game with character personalities and goals that are influenced by social norms. Roberto Arias-Yacupoma, Louis Katchen, Mubbasir Kapadia |
MIG | 4 |
| 2016 | An event-centric approach to authoring stories in crowdsabstractWe present a graphical authoring tool for creating complex narratives in large, populated areas with crowds of virtual humans. With an intuitive drag-and-drop interface, our system enables an untrained author to assemble story arcs in terms of narrative events that seamlessly control either principal characters or choreographed heterogeneous crowds within the same conceptual structure. Smart Crowds allow groups of characters to be dynamically assembled and scheduled with ambient activities, while also permitting individual characters to be selected from the crowd and featured more prominently as an individual in a story with more sophisticated behavior. Our system runs in real-time at interactive rates with no pause or costly pre-computation step between creating a story and simulating it, making this approach ideal for storyboarding or pre-visualization of narrative sequences. Mubbasir Kapadia, Alexander Shoulson, Cyril Steimer, Samuel Oberholzer, Robert W. Sumner, Markus Gross 0001 |
MIG | 1 |
| 2016 | KINDLING: a game platform for crowd-sourcing fire evacuation dataabstractKINDLING sets forth to make an engaging and intuitive platform for fire evacuation analytics. By modeling the problem as a game, real or imaginary buildings can be created as game levels with a level editor, and attacked by other players, providing a highly advanced worst case scenario. Attacking players use a limited number of fires to attempt the most efficient destruction of the simulated crowd in a level designed by another player. Scores are recorded and displayed much like any other video game. However, each attempt at the level collects data for important evacuation metrics. With this data represented as a heat-map overlay on the level, the creator can improve upon the level by adding precautionary measures to dangerous locations; such as an extinguisher to clear the way to an exit, or a fire door to hold back the spread of the fire. In this way, KINDLING provides both meaningful feedback on the evacuation safety of a real or virtual space, and evolving dynamic game-play between the level creator and other players. Leonard Wohl, Roland Gorzkowski, Nicole Fox, Mubbasir Kapadia |
MIG | 4 |
| 2016 | Precision: precomputing environment semantics for contact-rich character animationabstractThe widespread availability of high-quality motion capture data and the maturity of solutions to animate virtual characters has paved the way for the next generation of interactive virtual worlds exhibiting intricate interactions between characters and the environments they inhabit. However, current motion synthesis techniques have not been designed to scale with complex environments and contact-rich motions, requiring environment designers to manually embed motion semantics in the environment geometry in order to address online motion synthesis. This paper presents an automated approach for analyzing both motions and environments in order to represent the different ways in which an environment can afford a character to move. We extract the salient features that characterize the contact-rich motion repertoire of a character and detect valid transitions in the environment where each of these motions may be possible, along with additional semantics that inform which surfaces of the environment the character may use for support during the motion. The precomputed motion semantics can be easily integrated into standard navigation and animation pipelines in order to greatly enhance the motion capabilities of virtual characters. The computational efficiency of our approach enables two additional applications. Environment designers can interactively design new environments and get instant feedback on how characters may potentially interact, which can be used for iterative modeling and refinement. End users can dynamically edit virtual worlds and characters will automatically accommodate the changes in the environment in their movement strategies. Mubbasir Kapadia, Xianghao Xu, Maurizio Nitti, Marcelo Kallmann, Stelian Coros, Robert W. Sumner, Markus Gross 0001 |
I3D | 1 |
| 2016 | ACCLMesh: curvature-based navigation mesh generationabstractAbstract The proposed method computes a navigation mesh for arbitrary and dynamic 3D environments based on curvature and is robust and efficient. This method addresses a number of known limitations in state‐of‐the‐art techniques to produce navigation meshes that are tightly coupled to the original geometry, incorporate geometric details that are crucial for movement decisions, can robustly handle complex surfaces and can efficiently repair the navigation mesh to accommodate dynamically changing environments. The method is integrated into a standard navigation and collision avoidance system to simulate thousands of agents on complex 3D surfaces in real time. Copyright © 2016 John Wiley & Sons, Ltd. Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 2 |
| 2015 | Evaluating the Authoring Complexity of Interactive Narratives for Augmented Reality Applications
Mubbasir Kapadia, Fabio Zünd, Jessica Falk, Marcel Marti, Robert W. Sumner |
FDG | 1 |
| 2015 | Statistical Analysis of Player Behavior in Minecraft
Stephan Müller 0002, Mubbasir Kapadia, Seth Frey, Severin Klingler, Richard P. Mann, Barbara Solenthaler, Robert W. Sumner, Markus Gross 0001 |
FDG | 2 |
| 2015 | HeapCraft: Understanding and Improving Player Collaboration in Minecraft
Stephan Müller 0002, Mubbasir Kapadia, Seth Frey, Severin Klingler, Richard P. Mann, Barbara Solenthaler, Robert W. Sumner, Markus Gross 0001 |
FDG | 2 |
| 2015 | Authoring Background Character Responses to Foreground Characters
Fernando Geraci, Mubbasir Kapadia |
ICIDS | 2 |
| 2015 | ACCLMesh: curvature-based navigation mesh generationabstractWe propose a method to robustly and efficiently compute a navigation mesh for arbitrary and dynamic 3D environments based on curvature. This method addresses a number of known limitations in state-of-the-art techniques to produce navigation meshes that are tightly coupled to the original geometry, incorporate geometric details that are crucial for movement decisions and robustly handle complex surfaces. We integrate the method into a standard navigation and collision-avoidance system to simulate thousands of agents on complex 3D surfaces in real-time. Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 2 |
| 2015 | Automated interactive narrative synthesis using dramatic theoryabstractCurrent systems for automatic narrative generation lack modularity in their authoring tools as well as the ability to accommodate a human player's interaction with the characters while simultaneously preserving narrative integrity. In this paper, we propose a logical formalism of a story that incorporates an exposition, rising action, climax, falling action, and a story resolution, conforming to the widely established and studied Freytag model of a narrative. Using computational tools, our system is able to automatically synthesize stories that are grounded in narrative theory. These logical representations of stories are transformed into its equivalent Parameterized Behavior Tree (PBT) representation to facilitate an animated discourse of the narrative by leveraging existing character animations tools. Next, we automatically transform these passive narratives into interactive narratives by introducing narrative revision and nudge -- two extensions that preserve narrative integrity while still allowing the player to assume control of any character at any point in the story, and the freedom to experience the story in any way he sees fit. Our results demonstrate the promise of leveraging computational intelligence for automated interactive narrative synthesis while being firmly established in classical narrative theory. Carlos Antonio Dominguez, Yuya Ichimura, Mubbasir Kapadia |
MIG | 3 |
| 2015 | Evaluating and optimizing level of service for crowd evacuationsabstractLevel of service (LoS) is a standard indicator, widely used in crowd management and urban design, for characterizing the service afforded by environments to crowds of specific densities. However, current LoS indicators are qualitative and rely on expert analysis. Computational approaches for crowd analysis and environment design require robust measures for characterizing the relationship between environments and crowd flow. M. Brandon Haworth, Muhammad Usman 0010, Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 4 |
| 2015 | HeapCraft: interactive data exploration and visualization tools for understanding and influencing player behavior in MinecraftabstractWe present HeapCraft: an open-source suite of interactive data exploration and visualization tools that allows researchers, server administrators and game designers to analyze and potentially influence player behavior in Minecraft. Our framework includes a telemetry system, several tools for visualizing and representing the collected data, and tools for modifying the game experience in controlled ways. Measures that we use to quantify and visualize player behavior and collaboration have been derived from a large data set containing 3451 player-hours from 908 players and 43 different servers. HeapCraft has been demonstrated on a variety of tasks including player behavior classification, as well as quantifying and improving collaboration of players on Minecraft servers. HeapCraft is freely available and serves to democratize game analytics for the Minecraft community at large. Stephan Müller 0002, Barbara Solenthaler, Mubbasir Kapadia, Seth Frey, Severin Klingler, Richard P. Mann, Robert W. Sumner, Markus Gross 0001 |
MIG | 3 |
| 2015 | Computer-assisted authoring of interactive narrativesabstractThis paper explores new authoring paradigms and computer-assisted authoring tools for free-form interactive narratives. We present a new design formalism, Interactive Behavior Trees (IBT's), which decouples the monitoring of user input, the narrative, and how the user may influence the story outcome. We introduce automation tools for IBT's, to help the author detect and automatically resolve inconsistencies in the authored narrative, or conflicting user interactions that may hinder story progression. We compare IBT's to traditional story graph representations and show that our formalism better scales with the number of story arcs, and the degree and granularity of user input. The authoring time is further reduced with the help of automation, and errors are completely avoided. Our approach enables content creators to easily author complex, branching narratives with multiple story arcs in a modular, extensible fashion while empowering players with the agency to freely interact with the characters in the story and the world they inhabit. Mubbasir Kapadia, Jessica Falk, Fabio Zünd, Marcel Marti, Robert W. Sumner, Markus Gross 0001 |
I3D | 1 |
| 2015 | Footstep parameterized motion blending using barycentric coordinates
Alejandro Beacco, Nuria Pelechano, Mubbasir Kapadia, Norman I. Badler |
Comput. Graph. | 3 |
| 2015 | Environment optimization for crowd evacuationabstractAbstract The layout of a building, real or virtual, affects the flow patterns of its intended users. It is well established, for example, that the placement of pillars at proper locations can often facilitate pedestrian flow during the evacuation of a building. Such considerations are therefore important for architects, game level developers, and others whose domains involve agents navigating through buildings. In this paper, we take the first steps towards developing a simulation framework that can be used to study the optimal placement of architectural elements, such as pillars or doors, for the purposes of facilitating dense pedestrian flow during the evacuation of a building. In particular, we show that the steering algorithms used to model the local navigation abilities of the agents significantly affect the results, which motivates the need for a statistically valid approach and further study. Copyright © 2015 John Wiley & Sons, Ltd. Glen Berseth, Muhammad Usman 0010, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 4 |
| 2015 | Generating a multiplicity of policies for agent steering in crowd simulationabstractAbstract Pedestrian steering algorithms range from completely procedural to entirely data‐driven, but the former grossly generalize across possible human behaviors and suffer computationally, whereas the latter are limited by the burden of ever‐increasing data samples. Our approach seeks the balanced middle ground by deriving a collection of machine‐learned policies based on the behavior of a procedural steering algorithm through the decomposition of the space of possible steering scenarios into steering contexts. The resulting algorithm scales well in the number of contexts, the use of new data sets to create new policies, and in the number of controlled agents as the policies become a simple evaluation of the rules asserted by the machine‐learning process. We also explore the use of synthetic data from an “oracle algorithm” that serves as an as‐needed source of samples, which can be stochastically polled for effective coverage. We observe that our approach produces pedestrian steering similar to that of the oracle steering algorithm, but with a significant performance boost. Runtime was reduced from hours under the oracle algorithm with 10 agents to on the order of 10 frames per second (FPS) with 3000 agents. We also analyze the nature of collisions in such a framework with no explicit collision avoidance. Copyright © 2014 John Wiley & Sons, Ltd. Cory D. Boatright, Mubbasir Kapadia, Jennie M. Shapira, Norman I. Badler |
Comput. Animat. Virtual Worlds | 2 |
| 2015 | Planning approaches to constraint-aware navigation in dynamic environmentsabstractAbstract Path planning is a fundamental problem in many areas, ranging from robotics and artificial intelligence to computer graphics and animation. Although there is extensive literature for computing optimal, collision‐free paths, there is relatively little work that explores the satisfaction of spatial constraints between objects and agents at the global navigation layer. This paper presents a planning framework that satisfies multiple spatial constraints imposed on the path. The type of constraints specified can include staying behind a building, walking along walls, or avoiding the line of sight of patrolling agents. We introduce two hybrid environment representations that balance computational efficiency and search space density to provide a minimal, yet sufficient, discretization of the search graph for constraint‐aware navigation. An extended anytime dynamic planner is used to compute constraint‐aware paths, while efficiently repairing solutions to account for varying dynamic constraints or an updating world model. We demonstrate the benefits of our method on challenging navigation problems in complex environments for dynamic agents using combinations of hard and soft, attracting and repelling constraints, defined by both static obstacles and moving obstacles. Copyright © 2014 John Wiley & Sons, Ltd. Kai Ninomiya, Mubbasir Kapadia, Alexander Shoulson, Francisco M. Garcia, Norman I. Badler |
Comput. Animat. Virtual Worlds | 2 |
| 2014 | GPU-based dynamic search on adaptive resolution gridsabstractThis paper presents a GPU-based wave-front propagation technique for multi-agent path planning in extremely large, complex, dynamic environments. Our work proposes an adaptive subdivision of the environment with efficient indexing, update, and neighbor-finding operations on the GPU to address several known limitations in prior work. In particular, an adaptive environment representation reduces the device memory requirements by an order of magnitude which enables for the first time, GPU-based goal path planning in truly large-scale environments (> 2048 m2) for hundreds of agents with different targets. We compare our approach to prior work that uses an uniform grid on several challenging navigation benchmarks and report significant memory savings, and up to a 1000X computational speedup. Francisco M. Garcia, Mubbasir Kapadia, Norman I. Badler |
ICRA | 2 |
| 2014 | Path planning for coherent and persistent groupsabstractThis paper addresses the problem of group path planning while maintaining group coherence and persistence. Group coherence ensures that a group minimizes both longitudinal and lateral dispersion, and is achieved with the introduction of a deformation penalty to the cost formulation. When the deformation penalty is significantly high, a group may split and later merge. Group persistence is modeled by introducing split and merge actions in the action space, and adding a split penalty to the cost measure. We formulate the problem domain (state, action space, and cost formulation), present our path planning approach for coherent and persistent groups, and provide empirical results demonstrating the capabilities of our method on a variety of challenging scenarios. Mubbasir Kapadia, Norman I. Badler, Marcelo Kallmann |
ICRA | 2 |
| 2014 | Characterizing and optimizing game level difficultyabstractBalancing the interactions between game level design and intended player experience is a difficult and time consuming process. Automating aspects of this process with respect to user-defined constraints has beneficial implications for game designers. A change in level layout may affect the available routes and subsequent player interactions for a number of agents within the level. Small changes in the placement of game elements may lead to significant changes in terms of the challenge experienced by the player on the path to their goal. Estimating the effect of this change requires that the designer take into account new paths of all interacting agents and how these may affect the player. As the number of these agents grow to crowd size, estimating the effect of these changes becomes grows difficult. We present a user-in-the-loop framework for tackling this task by optimizing enemy agent settings and the placement of game elements that affect the flow of agents within the level, with respect to estimated difficulty. Using static path analysis we estimate difficulty based on agent interactions with the player. To exemplify the usefulness of the framework, we show that small changes in level layout lead to significant changes in game difficulty, and optimizations with respect to the characterization of difficulty can be used to attain desired difficulty levels. Glen Berseth, M. Brandon Haworth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 3 |
| 2014 | Sound localization and multi-modal steering for autonomous virtual agentsabstractWith the increasing realism of interactive applications, there is a growing need for harnessing additional sensory modalities such as hearing. While the synthesis and propagation of sounds in virtual environments has been explored, there has been little work that addresses sound localization and its integration into behaviors for autonomous virtual agents. This paper develops a framework that enables autonomous virtual agents to localize sounds in dynamic virtual environments, subject to distortion effects due to attenuation, reflection and diffraction from obstacles, as well as interference between multiple audio signals. We additionally integrate hearing into standard predictive collision avoidance techniques and couple it with vision to allow agents to react to what they see and hear, while navigating in virtual environments. Yu Wang 0033, Mubbasir Kapadia, Ladislav Kavan, Norman I. Badler |
I3D | 2 |
| 2014 | ADAPT: The Agent Developmentand Prototyping TestbedabstractWe present ADAPT, a flexible platform for designing and authoring functional, purposeful human characters in a rich virtual environment. Our framework incorporates character animation, navigation, and behavior with modular interchangeable components to produce narrative scenes. The animation system provides locomotion, reaching, gaze tracking, gesturing, sitting, and reactions to external physical forces, and can easily be extended with more functionality due to a decoupled, modular structure. The navigation component allows characters to maneuver through a complex environment with predictive steering for dynamic obstacle avoidance. Finally, our behavior framework allows a user to fully leverage a character's animation and navigation capabilities when authoring both individual decision-making and complex interactions between actors using a centralized, event-driven model. Alexander Shoulson, Nathan Marshak, Mubbasir Kapadia, Norman I. Badler |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2013 | The effect of posture and dynamics on the perception of emotionabstractMotion capture remains a popular and widely-used method for animating virtual characters. However, all practical applications of motion capture rely on motion editing techniques to increase the reusability and flexibility of captured motions. Because humans are proficient in detecting and interpreting subtle details in human motion, understanding the perceptual consequences of motion editing is essential. Thus in this work, we perform three experiments to gain a better understanding of how motion editing might affect the emotional content of a captured performance, particularly changes in posture and dynamics, two factors shown to be important perceptual indicators of bodily emotions. In these studies, we analyse the properties (angles and velocities) and perception (recognition rates and perceived intensities) of a varied set of full-body motion clips representing the six emotions anger, disgust, fear, happiness, sadness, and surprise. We have found that emotions are mostly conveyed through the upper body, that the perceived intensity of an emotion can be reduced by blending with a neutral motion, and that posture changes can alter the perceived emotion but subtle changes in dynamics only alter the intensity. Aline Normoyle, Fannie Liu, Mubbasir Kapadia, Norman I. Badler, Sophie Jörg |
SAP | 3 |
| 2013 | Dynamic search on the GPUabstractPath finding is a fundamental, yet computationally expensive problem in robotics navigation. Often times, it is necessary to sacrifice optimality to find a feasible plan given a time constraint due to the search complexity. Dynamic environments may further invalidate current computed plans, requiring an efficient planning strategy that can repair existing solutions. This paper presents a massively parallelized wavefront-based approach to path planning, running on the GPU, that can efficiently repair plans to accommodate world changes and agent movement, without having to restart the wavefront propagation process. In addition, we introduce a termination condition which ensures the minimum number of GPU iterations while maintaining strict optimality constraints on search graphs with non-uniform costs. Mubbasir Kapadia, Francisco M. Garcia, Cory D. Boatright, Norman I. Badler |
IROS | 1 |
| 2013 | SteerPlex: Estimating Scenario Complexity for Simulated CrowdsabstractThe complexity of interactive virtual worlds has increased dramatically in recent years, with a rise in mature solutions for designing large-scale environments and populating them with hundreds and thousands of autonomous characters. An interesting problem that arises in this context, and that has received little attention to date, is whether we can predict the complexity of a steering scenario by analyzing the configuration of the environment and the agents involved. We statically analyze an input scenario and compute a set of novel salient features which characterize the expected interactions between agents and obstacles during simulation. Using a statistical approach, we automatically derive the relative influence of each feature on the complexity of a scenario in order to derive a single numerical quantity of expected scenario complexity. We validate our proposed metric by demonstrating a strong negative correlation between the statically computed expected complexity and the dynamic performance of three published crowd simulation techniques. Glen Berseth, Mubbasir Kapadia, Petros Faloutsos |
MIG | 2 |
| 2013 | Constraint-Aware Navigation in Dynamic EnvironmentsabstractPath planning is a fundamental problem in many areas ranging from robotics and artificial intelligence to computer graphics and animation. While there is extensive literature for computing optimal, collision-free paths, there is little work that explores the satisfaction of spatial constraints between objects and agents at the global navigation layer. This paper presents a planning framework that satisfies multiple spatial constraints imposed on the path. The type of constraints specified could include staying behind a building, walking along walls, or avoiding the line of sight of patrolling agents. We introduce a hybrid environment representation that balances computational efficiency and discretization resolution, to provide a minimal, yet sufficient discretization of the search graph for constraint-aware navigation. An extended anytime-dynamic planner is used to compute constraint-aware paths, while efficiently repairing solutions to account for dynamic constraints. We demonstrate the benefits of our method on challenging navigation problems in complex environments for dynamic agents using combinations of hard and soft constraints, attracting and repelling constraints, on static obstacles and moving obstacles. Mubbasir Kapadia, Kai Ninomiya, Alexander Shoulson, Francisco M. Garcia, Norman I. Badler |
MIG | 1 |
| 2013 | An Event-Centric Planning Approach for Dynamic Real-Time NarrativeabstractIn this paper, we propose an event-centric planning framework for directing interactive narratives in complex 3D environments populated by virtual humans. Events facilitate precise authorial control over complex interactions involving groups of actors and objects, while planning allows the simulation of causally consistent character actions that conform to an overarching global narrative. Events are defined by preconditions, postconditions, costs, and a centralized behavior structure that simultaneously manages multiple participating actors and objects. By planning in the space of events rather than in the space of individual character capabilities, we allow virtual actors to exhibit a rich repertoire of individual actions without causing combinatorial growth in the planning branching factor. Our system produces long, cohesive narratives at interactive rates, allowing a user to take part in a dynamic story that, despite intervention, conforms to an authored structure and accomplishes a predetermined goal. Alexander Shoulson, Max L. Gilbert, Mubbasir Kapadia, Norman I. Badler |
MIG | 3 |
| 2013 | Efficient motion retrieval in large motion databasesabstractThere has been a recent paradigm shift in the computer animation industry with an increasing use of pre-recorded motion for animating virtual characters. A fundamental requirement to using motion capture data is an efficient method for indexing and retrieving motions. In this paper, we propose a flexible, efficient method for searching arbitrarily complex motions in large motion databases. Motions are encoded using keys which represent a wide array of structural, geometric and, dynamic features of human motion. Keys provide a representative search space for indexing motions and users can specify sequences of key values as well as multiple combination of key sequences to search for complex motions. We use a trie-based data structure to provide an efficient mapping from key sequences to motions. The search times (even on a single CPU) are very fast, opening the possibility of using large motion data sets in real-time applications. Mubbasir Kapadia, I-Kao Chiang, Tiju Thomas, Norman I. Badler, Joseph T. Kider Jr. |
I3D | 1 |
| 2013 | ADAPT: the agent development and prototyping testbedabstractWe present ADAPT, a flexible platform for designing and authoring functional, purposeful human characters in a rich virtual environment. Our framework incorporates character animation, navigation, and behavior with modular interchangeable components to produce narrative scenes. Our animation system provides locomotion, reaching, gaze tracking, gesturing, sitting, and reactions to external physical forces, and can easily be extended with more functionality due to a decoupled, modular structure. Additionally, our navigation component allows characters to maneuver through a complex environment with predictive steering for dynamic obstacle avoidance. Finally, our behavior framework allows a user to fully leverage a character's animation and navigation capabilities when authoring both individual decision-making and complex interactions between actors using a centralized, event-driven model. Alexander Shoulson, Nathan Marshak, Mubbasir Kapadia, Norman I. Badler |
I3D | 3 |
| 2012 | What's Next? The New Era of Autonomous Virtual Humans
Mubbasir Kapadia, Alexander Shoulson, Cory D. Boatright, Funda Durupinar, Norman I. Badler |
MIG | 1 |
| 2012 | Parallelized egocentric fields for autonomous navigation
Mubbasir Kapadia, Shawn Singh, William Hewlett, Glenn Reinman, Petros Faloutsos |
Vis. Comput. | 1 |
| 2011 | Improved Benchmarking for Steering Algorithms
Mubbasir Kapadia, Matthew Wang, Glenn Reinman, Petros Faloutsos |
MIG | 1 |
| 2011 | Parallelized Incomplete Poisson Preconditioner in Cloth Simulation
Costas Sideris, Mubbasir Kapadia, Petros Faloutsos |
MIG | 2 |
| 2011 | Behavior authoring for crowd simulationsabstractThere has been growing academic and industry interest in the behavioral animation of autonomous actors in virtual worlds. However, it remains a considerable challenge to automatically generate complicated interactions between multiple actors in a customizable way with minimal user specification. Mubbasir Kapadia, Shawn Singh, Glenn Reinman, Petros Faloutsos |
SI3D | 1 |
| 2011 | A modular framework for adaptive agent-based steeringabstractNext-generation steering algorithms will need to support thousands of believable individual agents, capable of steering in very challenging situations with low-latency reactions. In this paper we propose a steering framework that offers three key contributions: (a) It integrates several models of steering into a single steering decision, (b) it employs a novel space-time planning approach to allow agents to steer during complex local interactions, and (c) it varies the frequency of update of each component (phase) of the framework to drastically improve performance. We demonstrate the versatility and robustness of our framework using a large number of test cases. We also show that the frequency of updates for each phase of the framework can be "decimated" by a surprisingly large amount before resulting steering behaviors degrade. This technique achieves more than a 5x performance improvement, allowing the use of better, more costly algorithms for robust steering, while supporting thousands of agents with low-latency reactions in real-time. Shawn Singh, Mubbasir Kapadia, William Hewlett, Glenn Reinman, Petros Faloutsos |
SI3D | 2 |
| 2011 | Footstep navigation for dynamic crowdsabstractThe majority of previous crowd 'steering algorithms model each character as an oriented particle that moves by choosing a force or velocity vector. In many cases, orientation is heuristically chosen to be the same as the particle's velocity. This approach has the two key disadvantages: Shawn Singh, Mubbasir Kapadia, Glenn Reinman, Petros Faloutsos |
SI3D | 2 |
| 2011 | Footstep navigation for dynamic crowdsabstractAbstract The majority of steering algorithms output only a force or velocity vector to an animation system, without modeling the constraints and capabilities of human‐like movement. This simplistic approach lacks control over how a character should navigate. This paper proposes a steering method that usesfootstepsto navigate characters in dynamic crowds. Instead of an oriented particle with a single collision radius, we model a character's center of mass and footsteps using a 2D approximation of an inverted spherical pendulum model of bipedal locomotion. We use this model to generate a timed sequence of footsteps that existing animation techniques can follow exactly. Our approach not only constrains characters to navigate with realistic steps but also enables characters to intelligently control subtlenavigationbehaviors that are possible with exact footsteps, such as side‐stepping. Our approach can navigate crowds of hundreds of individual characters with collision‐free, natural steering decisions in real‐time. Copyright © 2011 John Wiley & Sons, Ltd. Shawn Singh, Mubbasir Kapadia, Glenn Reinman, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 2 |
| 2010 | Full-Body Hybrid Motor Control for Reaching
Wenjia Huang, Mubbasir Kapadia, Demetri Terzopoulos |
MIG | 2 |
| 2010 | Situation agents: agent-based externalized steering logicabstractAbstract We present a simple and intuitive method for encapsulating part of agents' steering and coordinating abilities into a new class of agents, called situation agents. Situation agents have all the abilities of typical agents. In addition, they can influence the steering decisions of any agent, including other situation agents, within their sphere of influence. Encapsulating steering logic into moving agents is a powerful abstraction which provides more flexibility and efficiency than traditional informed environment approaches, and works with many of the current steering methodologies. We demonstrate our proposed approach in a number of challenging scenarios. Copyright © 2010 John Wiley & Sons, Ltd. Matthew Schuerman, Shawn Singh, Mubbasir Kapadia, Petros Faloutsos |
Comput. Animat. Virtual Worlds | 3 |
| 2009 | Egocentric affordance fields in pedestrian steeringabstractIn this paper we propose a general framework for local path-planning and steering that can be easily extended to perform high-level behaviors. Our framework is based on the concept of affordances - the possible ways an agent can interact with its environment. Each agent perceives the environment through a set of vector and scalar fields that are represented in the agent's local space. This egocentric property allows us to efficiently compute a local space-time plan. We then use these perception fields to compute a fitness measure for every possible action, known as an affordance field. The action that has the optimal value in the affordance field is the agent's steering decision. Using our framework, we demonstrate autonomous virtual pedestrians that perform steering and path planning in unknown environments along with the emergence of high-level responses to never seen before situations. Mubbasir Kapadia, Shawn Singh, William Hewlett, Petros Faloutsos |
SI3D | 1 |
| 2009 | SteerBench: a benchmark suite for evaluating steering behaviorsabstractAbstract Steering is a challenging task, required by nearly all agents in virtual worlds. There is a large and growing number of approaches for steering, and it is becoming increasingly important to ask a fundamental question: how can we objectively compare steering algorithms? To our knowledge, there is no standard way of evaluating or comparing the quality of steering solutions. This paper presents SteerBench: a benchmark framework for objectively evaluating steering behaviors for virtual agents. We propose a diverse set of test cases, metrics of evaluation, and a scoring method that can be used to compare different steering algorithms. Our framework can be easily customized by a user to evaluate specific behaviors and new test cases. We demonstrate our benchmark process on two example steering algorithms, showing the insight gained from our metrics. We hope that this framework can grow into a standard for steering evaluation. Copyright © 2009 John Wiley & Sons, Ltd. Shawn Singh, Mubbasir Kapadia, Petros Faloutsos, Glenn Reinman |
Comput. Animat. Virtual Worlds | 2 |