EDBT 2026 Demo / reviewers in the wild / expert
Samuel S. Sohn
dblp:223/0144
· DBLP profile ↗
28ranked-venue papers
8as first author
20since 2021 · last 2025
0000-0003-4700-954XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Separability of Human Navigational Behaviors in Virtual Reality
Samuel S. Sohn, Serena DeStefani, Mathew Schwartz, Jacob Feldman, Mubbasir Kapadia, Karin Stromswold |
CogSci | 1 |
| 2025 | Prosody in the Age of AI: Insights from Large Speech Models
Samuel S. Sohn, Sten Knutsen, Karin Stromswold |
CogSci | 1 |
| 2025 | Cardiverse: Harnessing LLMs for Novel Card Game PrototypingabstractThe prototyping of computer games, particularly card games, requires extensive human effort in creative ideation and gameplay evaluation.Recent advances in Large Language Models (LLMs) offer opportunities to automate and streamline these processes.However, it remains challenging for LLMs to design novel game mechanics beyond existing databases, generate consistent gameplay environments, and develop scalable gameplay AI for large-scale evaluations.This paper addresses these challenges by introducing a comprehensive automated card game prototyping framework.The approach highlights a graph-based indexing method for generating novel game variations, an LLM-driven system for consistent game code generation validated by gameplay records, and a gameplay AI constructing method that uses an ensemble of LLM-generated heuristic functions optimized through self-play.These contributions aim to accelerate card game prototyping, reduce human labor, and lower barriers to entry for game developers. Danrui Li, Samuel S. Sohn, Kaidong Hu, Muhammad Usman 0010, Mubbasir Kapadia |
EMNLP | 3 |
| 2024 | Learning from Synthetic Human Group ActivitiesabstractThe study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation, we introduce M3 Act, a synthetic data generator for multi-view multi-group multi-person human atomic actions and group activities. Powered by Unity Engine, M3 Act features mul-tiple semantic groups, highly diverse and photorealistic images, and a comprehensive set of annotations, which facilitates the learning of human-centered tasks across single-person, multi-person, and multi-group conditions. We demonstrate the advantages of M3 Act across three core experiments. The results suggest our synthetic dataset can significantly improve the performance of several downstream methods and replace real-world datasets to reduce cost. Notably, M3 Act improves the state-of-the-art MOTRv2 on DanceTrack dataset, leading to a hop on the leaderboard from 10thto 2ndplace. Moreover, M3 Act opens new research for controllable 3D group activity generation. We define multiple metrics and propose a competitive baseline for the novel task. Our code and data are available at our project page: http://cjerry1243.github.io/M3Act. Che-Jui Chang, Danrui Li, Deep Patel, Parth Goel, Honglu Zhou, Seonghyeon Moon, Samuel S. Sohn, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
CVPR | 7 |
| 2024 | TrajDiffuse: A Conditional Diffusion Model for Environment-Aware Trajectory Prediction
Tony Qingze Liu, Danrui Li, Samuel S. Sohn, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001 |
ICPR (29) | 3 |
| 2024 | From Words to Worlds: Transforming One-line Prompts into Multi-modal Digital Stories with LLM AgentsabstractDigital storytelling, essential in entertainment, education, and marketing, faces challenges in generation efficiency. The StoryAgent framework, introduced in this paper, utilizes Large Language Models and generative tools to automate and refine digital storytelling. Employing a top-down story drafting and bottom-up asset generation approach, StoryAgent tackles key issues such as manual intervention, interactive scene orchestration, and narrative consistency. This framework enables efficient production of interactive and consistent digital storytellings across multiple modalities, democratizing content creation and enhancing engagement. Danrui Li, Samuel S. Sohn, Che-Jui Chang, Mubbasir Kapadia |
MIG | 2 |
| 2023 | MSI: Maximize Support-Set Information for Few-Shot SegmentationabstractFSS (Few-shot segmentation) aims to segment a target class using a small number of labeled images (support set). To extract information relevant to the target class, a dominant approach in best performing FSS methods removes background features using a support mask. We observe that this feature excision through a limiting support mask introduces an information bottleneck in several challenging FSS cases, e.g., for small targets and/or inaccurate target boundaries. To this end, we present a novel method (MSI), which maximizes the support-set information by exploiting two complementary sources of features to generate super correlation maps. We validate the effectiveness of our approach by instantiating it into three recent and strong FSS methods. Experimental results on several publicly available FSS benchmarks show that our proposed method consistently improves performance by visible margins and leads to faster convergence. Our code and trained models are available at: https://github.com/moonsh/MSI-Maximize-Support-Set-Information Seonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon, Vladimir Pavlovic 0001, Muhammad Haris Khan, Mubbasir Kapadia |
ICCV | 2 |
| 2023 | Harnessing Neighborhood Modeling and Asymmetry Preservation for Digraph Representation LearningabstractDigraph Representation Learning aims to learn representations for directed homogeneous graphs (digraphs). Prior work is largely constrained or has poor generalizability across tasks. Most Graph Neural Networks exhibit poor performance on digraphs due to the neglect of modeling neighborhoods and preserving asymmetry. In this paper, we address these notable challenges by leveraging hyperbolic collaborative learning from multi-ordered partitioned neighborhoods and asymmetry-preserving regularizers. Our resulting formalism, Digraph Hyperbolic Networks (D-HYPR), is versatile for multiple tasks including node classification, link presence prediction, and link property prediction. The efficacy of D-HYPR was meticulously examined against 21 previous techniques, using 8 real-world digraph datasets. D-HYPR statistically significantly outperforms the current state of the art. We release our code at https://github. com/hongluzhou/dhypr. Honglu Zhou, Advith Chegu, Samuel S. Sohn, Zuohui Fu, Gerard de Melo, Mubbasir Kapadia |
IJCAI | 3 |
| 2023 | The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational AgentsabstractPrevious studies regarding the perception of emotions for embodied virtual agents have shown the effectiveness of using virtual characters in conveying emotions through interactions with humans. However, creating an autonomous embodied conversational agent with expressive behaviors presents two major challenges. The first challenge is the difficulty of synthesizing the conversational behaviors for each modality that are as expressive as real human behaviors. The second challenge is that the affects are modeled independently, which makes it difficult to generate multimodal responses with consistent emotions across all modalities. In this work, we propose a conceptual framework, ACTOR (Affect-Consistent mulTimodal behaviOR generation), that aims to increase the perception of affects by generating multimodal behaviors conditioned on a consistent driving affect. We have conducted a user study with 199 participants to assess how the average person judges the affects perceived from multimodal behaviors that are consistent and inconsistent with respect to a driving affect. The result shows that among all model conditions, our affect-consistent framework receives the highest Likert scores for the perception of driving affects. Our statistical analysis suggests that making a modality affect-inconsistent significantly decreases the perception of driving affects. We also observe that multimodal behaviors conditioned on consistent affects are more expressive compared to behaviors with inconsistent affects. Therefore, we conclude that multimodal emotion conditioning and affect consistency are vital to enhancing the perception of affects for embodied conversational agents. Che-Jui Chang, Samuel S. Sohn, Rajath Jayashankar, Muhammad Usman 0010, Mubbasir Kapadia |
IUI | 2 |
| 2023 | Cognitive Path Planning With Spatial Memory DistortionabstractHuman path-planning operates differently from deterministic AI-based path-planning algorithms due to the decay and distortion in a human's spatial memory and the lack of complete scene knowledge. Here, we present a cognitive model of path-planning that simulates human-like learning of unfamiliar environments, supports systematic degradation in spatial memory, and distorts spatial recall during path-planning. We propose a Dynamic Hierarchical Cognitive Graph (DHCG) representation to encode the environment structure by incorporating two critical spatial memory biases during exploration: categorical adjustment and sequence order effect. We then extend the "Fine-To-Coarse" (FTC), the most prevalent path-planning heuristic, to incorporate spatial uncertainty during recall through the DHCG. We conducted a lab-based Virtual Reality (VR) experiment to validate the proposed cognitive path-planning model and made three observations: (1) a statistically significant impact of sequence order effect on participants' route-choices, (2) approximately three hierarchical levels in the DHCG according to participants' recall data, and (3) similar trajectories and significantly similar wayfinding performances between participants and simulated cognitive agents on identical path-planning tasks. Furthermore, we performed two detailed simulation experiments with different FTC variants on a Manhattan-style grid. Experimental results demonstrate that the proposed cognitive path-planning model successfully produces human-like paths and can capture human wayfinding's complex and dynamic nature, which traditional AI-based path-planning algorithms cannot capture. Rohit Kumar Dubey, Samuel S. Sohn, Tyler Thrash, Christoph Hölscher, André Borrmann, Mubbasir Kapadia |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | D-HYPR: Harnessing Neighborhood Modeling and Asymmetry Preservation for Digraph Representation LearningabstractDigraph Representation Learning (DRL) aims to learn representations for directed homogeneous graphs (digraphs). Prior work in DRL is largely constrained (e.g., limited to directed acyclic graphs), or has poor generalizability across tasks (e.g., evaluated solely on one task). Most Graph Neural Networks (GNNs) exhibit poor performance on digraphs due to the neglect of modeling neighborhoods and preserving asymmetry. In this paper, we address these notable challenges by leveraging hyperbolic collaborative learning from multi-ordered and partitioned neighborhoods, and regularizers inspired by socio-psychological factors. Our resulting formalism, Digraph Hyperbolic Networks (D-HYPR) -- albeit conceptually simple -- generalizes to digraphs where cycles and non-transitive relations are common, and is applicable to multiple downstream tasks including node classification, link presence prediction, and link property prediction. In order to assess the effectiveness of D-HYPR, extensive evaluations were performed across 8 real-world digraph datasets involving 21 prior techniques. D-HYPR statistically significantly outperforms the current state of the art. We release our code at https://github.com/hongluzhou/dhypr Honglu Zhou, Advith Chegu, Samuel S. Sohn, Zuohui Fu, Gerard de Melo, Mubbasir Kapadia |
CIKM | 3 |
| 2022 | How Optimal is Too Optimal? Expectations About Performance in the Traveling Salesman Problem
Serena De Stefani, Samuel S. Sohn, Adeeb Kabir, Mubbasir Kapadia, Jacob Feldman |
CogSci | 2 |
| 2022 | A Computational Method for the Classification of Mental Representations of Objects in 3D Space (Short Paper)
Samuel S. Sohn, Panagiotis Mavros, Mubbasir Kapadia, Christoph Hölscher |
COSIT | 1 |
| 2022 | MUSE-VAE: Multi-Scale VAE for Environment-Aware Long Term Trajectory PredictionabstractAccurate long-term trajectory prediction in complex scenes, where multiple agents (e.g., pedestrians or vehicles) interact with each other and the environment while attempting to accomplish diverse and often unknown goals, is a challenging stochastic forecasting problem. In this work, we propose MUSEVAE, a new probabilistic modeling framework based on a cascade of Conditional VAEs, which tackles the long-term, uncertain trajectory prediction task using a coarse-to-fine multi-factor forecasting architecture. In its Macro stage, the model learns a joint pixel-space representation of two key factors, the underlying environment and the agent movements, to predict the long and short term motion goals. Conditioned on them, the Micro stage learns a fine-grained spatio-temporal representation for the prediction of individual agent trajectories. The VAE backbones across the two stages make it possible to naturally account for the joint uncertainty at both levels of granularity. As a result, MUSEVAE offers diverse and simultaneously more accurate predictions compared to the current state-of-the-art. We demonstrate these assertions through a comprehensive set of experiments on nuScenes and SDD benchmarks as well as PFSD, a new synthetic dataset, which challenges the forecasting ability of models on complex agent-environment interaction scenarios. Mihee Lee, Samuel S. Sohn, Seonghyeon Moon, Sejong Yoon, Mubbasir Kapadia, Vladimir Pavlovic 0001 |
CVPR | 2 |
| 2022 | HM: Hybrid Masking for Few-Shot Segmentation
Seonghyeon Moon, Samuel S. Sohn, Honglu Zhou, Sejong Yoon, Vladimir Pavlovic 0001, Muhammad Haris Khan, Mubbasir Kapadia |
ECCV (20) | 2 |
| 2022 | Harnessing Fourier Isovists and Geodesic Interaction for Long-Term Crowd Flow PredictionabstractWith the rise in popularity of short-term Human Trajectory Prediction (HTP), Long-Term Crowd Flow Prediction (LTCFP) has been proposed to forecast crowd movement in large and complex environments. However, the input representations, models, and datasets for LTCFP are currently limited. To this end, we propose Fourier Isovists, a novel input representation based on egocentric visibility, which consistently improves all existing models. We also propose GeoInteractNet (GINet), which couples the layers between a multi-scale attention network (M-SCAN) and a convolutional encoder-decoder network (CED). M-SCAN approximates a super-resolution map of where humans are likely to interact on the way to their goals and produces multi-scale attention maps. The CED then uses these maps in either its encoder's inputs or its decoder's attention gates, which allows GINet to produce super-resolution predictions with substantially higher accuracy than existing models even with Fourier Isovists. In order to evaluate the scalability of models to large and complex environments, which the only existing LTCFP dataset is unsuitable for, a new synthetic crowd dataset with both real and synthetic environments has been generated. In its nascent state, LTCFP has much to gain from our key contributions. The Supplementary Materials, dataset, and code are available at sssohn.github.io/GeoInteractNet. Samuel S. Sohn, Seonghyeon Moon, Honglu Zhou, Mihee Lee, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
IJCAI | 1 |
| 2022 | A2X: An end-to-end framework for assessing agent and environment interactions in multimodal human trajectory prediction
Samuel S. Sohn, Mihee Lee, Seonghyeon Moon, Gang Qiao, Muhammad Usman 0010, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
Comput. Graph. | 1 |
| 2021 | The Interplay Between Local and Global Strategies in Navigational Decisions
Serena De Stefani, Samuel S. Sohn, Adeeb Kabir, Mubbasir Kapadia, Jacob Feldman |
CogSci | 2 |
| 2021 | SNAP: Successor Entropy based Incremental Subgoal Discovery for Adaptive NavigationabstractReinforcement learning (RL) has demonstrated great success in solving navigation tasks but often fails when learning complex environmental structures. One open challenge is to incorporate low-level generalizable skills with human-like adaptive path-planning in an RL framework. Motivated by neural findings in animal navigation, we propose a Successor eNtropy-based Adaptive Path-planning (SNAP) that combines a low-level goal-conditioned policy with the flexibility of a classical high-level planner. SNAP decomposes distant goal-reaching tasks into multiple nearby goal-reaching sub-tasks using a topological graph. To construct this graph, we propose an incremental subgoal discovery method that leverages the highest-entropy states in the learned Successor Representation. The Successor Representation encodes the likelihood of being in a future state given the current state and capture the relational structure of states based on a policy. Our main contributions lie in discovering subgoal states that efficiently abstract the state-space and proposing a low-level goal-conditioned controller for local navigation. Since the basic low-level skill is learned independent of state representation, our model easily generalizes to novel environments without intensive relearning. We provide empirical evidence that the proposed method enables agents to perform long-horizon sparse reward tasks quickly, take detours during barrier tasks, and exploit shortcuts that did not exist during training. Our experiments further show that the proposed method outperforms the existing goal-conditioned RL algorithms in successfully reaching distant-goal tasks and policy learning. To evaluate human-like adaptive path-planning, we also compare our optimal agent with human data and found that, on average, the agent was able to find a shorter path than the human participants. Rohit Kumar Dubey, Samuel S. Sohn, Jimmy Abualdenien, Tyler Thrash, Christoph Hölscher, André Borrmann, Mubbasir Kapadia |
MIG | 2 |
| 2021 | A2X: An Agent and Environment Interaction Benchmark for Multimodal Human Trajectory PredictionabstractIn recent years, human trajectory prediction (HTP) has garnered attention in computer vision literature. Although this task has much in common with the longstanding task of crowd simulation, there is little from crowd simulation that has been borrowed, especially in terms of evaluation protocols. The key difference between the two tasks is that HTP is concerned with forecasting multiple steps at a time and capturing the multimodality of real human trajectories. A majority of HTP models are trained on the same few datasets, which feature small, transient interactions between real people and little to no interaction between people and the environment. Unsurprisingly, when tested on crowd egress scenarios, these models produce erroneous trajectories that accelerate too quickly and collide too frequently, but the metrics used in HTP literature cannot convey these particular issues. To address these challenges, we propose (1) the A2X dataset, which has simulated crowd egress and complex navigation scenarios that compensate for the lack of agent-to-environment interaction in existing real datasets, and (2) evaluation metrics that convey model performance with more reliability and nuance. A subset of these metrics are novel multiverse metrics, which are better-suited for multimodal models than existing metrics. The dataset is available at: https://mubbasir.github.io/HTP-benchmark/. Samuel S. Sohn, Mihee Lee, Seonghyeon Moon, Gang Qiao, Muhammad Usman 0010, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
MIG | 1 |
| 2020 | Laying the Foundations of Deep Long-Term Crowd Flow Prediction
Samuel S. Sohn, Honglu Zhou, Seonghyeon Moon, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia |
ECCV (29) | 1 |
| 2019 | Fusion-Based Wayfinding Prediction Model for Multiple Information Sources
Rohit Kumar Dubey, Samuel S. Sohn, Christoph Hölscher, Mubbasir Kapadia |
FUSION | 2 |
| 2019 | JUNGLE: An Interactive Visual Platform for Collaborative Creation and Consumption of Nonlinear Transmedia Stories
Mubbasir Kapadia, Carlos Muñiz 0001, Samuel S. Sohn, Sasha Schriber, Kenny Mitchell, Markus Gross 0001 |
ICIDS | 3 |
| 2019 | StoryPrint: an interactive visualization of storiesabstractIn this paper, we propose StoryPrint, an interactive visualization of creative storytelling that facilitates individual and comparative structural analyses. This visualization method is intended for script-based media, which has suitable metadata. The pre-visualization process involves parsing the script into different metadata categories and analyzing the sentiment on a character and scene basis. For each scene, the setting, character presence, character prominence, and character emotion of a film are represented as a StoryPrint. The visualization is presented as a radial diagram of concentric rings wrapped around a circular time axis. A user then has the ability to toggle a difference overlay to assist in the cross-comparison of two different scene inputs. Katie Watson, Samuel S. Sohn, Sasha Schriber, Markus Gross 0001, Carlos Muñiz 0001, Mubbasir Kapadia |
IUI | 2 |
| 2019 | Towards a Conversational Interface for Authoring Intelligent Virtual CharactersabstractThe collaboration between creatives and domain architects is crucial for bringing virtual characters to life. Domain architects are technical experts who are tasked with formally designing intelligent virtual characters' domain knowledge, which is a symbolic representation of knowledge that the character uses to reason over its interactions with other agents. In the context of this work, domain knowledge encompasses the mental modeling of the character. Although the creation of interactive narratives requires substantial engineering expertise, it is also necessary to pick the brains of writers, artists, and animators alike to give the characters a boost of peculiarities. This intrinsically collaborative and interdisciplinary process brings about the challenge of bridging different mindsets and workflows in an efficient and effective way. Samuel S. Sohn, Mubbasir Kapadia |
IVA | 2 |
| 2019 | Identifying Indoor Navigation Landmarks Using a Hierarchical Multi-Criteria Decision FrameworkabstractLandmarks play a vital role in human wayfinding by providing the structure for mental spatial representations and indicating locations with which to orient. Less research effort has been allocated towards automated landmark identification in indoor environments despite a growing interest in indoor navigation in the scientific community. In this paper, we propose a computational framework to identify indoor landmarks that is based on a hierarchical multi-criteria decision model and grounded in theories of spatial cognition and human information processing. Our model of landmark salience is represented as a hierarchical integration process of low-level features derived from a three-part, higher-level, salience vector (i.e., cognitive, spatial, and subjective salience). We use a fuzzy hierarchical composite-weighted (objective and subjective) Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) to derive the rankings for identified objects at decision points (i.e., intersections). The top N objects are then selected and compared to a list of landmarks derived from an eye-tracking based virtual reality (VR) experiment. A substantial overlap of 79% was observed between these two lists. The proposed framework is capable of reliably and accurately detecting indoor landmarks, which can be employed in the development of landmark-based robot/autonomous agent motion and indoor guidance systems. Rohit Kumar Dubey, Samuel S. Sohn, Tyler Thrash, Christoph Hölscher, Mubbasir Kapadia |
MIG | 2 |
| 2018 | Efficiency in Solving the Traveling Salesman Problem as Predictor of Perceived Humanness
Serena De Stefani, Samuel S. Sohn, Jacob Feldman, Mubbasir Kapadia, Peter C. Pantelis |
CogSci | 2 |
| 2018 | Dynamic cognitive maps for agent landmark navigation in unseen environmentsabstractThe development of autonomous agents for wayfinding tasks has long maintained the usage of naive, omniscient models for navigation. The simplicity of these models improves the scalability of crowd simulations, but limits the utility of such simulations to the visualization of general behaviors. This restricted scope does not allow for the observation of more nuanced, individualized behaviors. In this paper, we demonstrate a novel framework for agent simulations that does not rely on omniscience. Instead, each agent is equipped with a memory architecture that enables wayfinding by maintaining a cognitive map of the space explored by the agent. Based on findings from simulation studies, cognitive science, and psychology, we describe a wayfinding procedure that simulates human behavior and human cognitive processes, incorporating landmark navigation, path integration, and memory. This cognitive approach makes observations of agent behavior more comparable to those of human behavior. Samuel S. Sohn, Serena De Stefani, Mubbasir Kapadia |
MIG | 1 |