VLDB 2026 Research / reviewers in the wild / expert
Shugao Ma
dblp:70/418
· DBLP profile ↗
32ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0002-4986-2221ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HuMoCon: Concept Discovery for Human Motion UnderstandingabstractWe present HuMoCon, a novel motion-video understanding framework designed for advanced human behavior analysis. The core of our method is a human motion concept discovery framework that efficiently trains multi-modal encoders to extract semantically meaningful and generalizable features. HuMoCon addresses key challenges in motion concept discovery for understanding and reasoning, including the lack of explicit multi-modality feature alignment and the loss of high-frequency information in masked autoencoding frameworks. Our approach integrates a feature alignment strategy that leverages video for contextual understanding and motion for fine-grained interaction modeling, further with a velocity reconstruction mechanism to enhance high-frequency feature expression and mitigate temporal over-smoothing. Comprehensive experiments on standard benchmarks demonstrate that HuMoCon enables effective motion concept discovery and significantly outperforms state-of-the-art methods in training large models for human motion understanding. We will open-source the associated code with our paper. Qihang Fang, Chengcheng Tang, Bugra Tekin, Shugao Ma, Yanchao Yang 0001 |
CVPR | 4 |
| 2025 | LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand DiffusionabstractCurrent research on generating 3D hand-object interaction motion primarily focuses on in-domain objects. Generalization to unseen objects is essential for practical applications, yet it remains both challenging and largely unexplored. In this paper, we propose LatentHOI, a novel approach designed to tackle the challenges of generalizing hand-object interaction synthesis to unseen objects. Our main insight lies in decoupling high-level temporal motion from fine-grained spatial hand-object interactions via a latent diffusion model coupled with a Grasping Variational Autoencoder (Grasp-VAE). This configuration introduces regularization by enforcing a conditional dependency between spatial grasping and temporal motion, as well as through the regularized latent space for better generalization ability. We conducted extensive experiments in an unseen-object setting on both single-hand grasping and bi-manual motion datasets, including GRAB, DexYCB†, and OakInk. Quantitative and qualitative evaluations demonstrate that our method significantly enhances the realism and physical plausibility of generated motions for unseen objects, both in single and bimanual manipulations, compared to the state-of-the-art. Muchen Li, Sammy Joe Christen, Chengde Wan, Yujun Cai, Renjie Liao 0001, Leonid Sigal, Shugao Ma |
CVPR | 7 |
| 2025 | Streaming Videollms for Real-Time Procedural Video Understanding
Dibyadip Chatterjee, Edoardo Remelli, Yale Song, Bugra Tekin, Abhay Mittal, Bharat Bhatnagar, Necati Cihan Camgöz, Shreyas Hampali, Eric Sauser, Shugao Ma, Angela Yao, Fadime Sener |
ICCV | 10 |
| 2025 | Spatial and temporal beliefs for mistake detection in assembly tasksabstractAssembly tasks, as an integral part of daily routines and activities, involve a series of sequential steps that are prone to error. This paper proposes a novel method for identifying ordering mistakes in assembly tasks based on knowledge-grounded beliefs. The beliefs comprise spatial and temporal aspects, each serving a unique role. Spatial beliefs capture the structural relationships among assembly components and indicate their topological feasibility. Temporal beliefs model the action preconditions and enforce sequencing constraints. Furthermore, we introduce a learning algorithm that dynamically updates and augments the belief sets online. To evaluate, we first test our approach in deducing predefined rules on synthetic data based on industry assembly. We also verify our approach on the real-world Assembly101 dataset, enhanced with annotations of component information. Our framework achieves superior performance in detecting ordering mistakes under both synthetic and real-world settings, highlighting the effectiveness of our approach. • We present two belief sets for the assembly tasks. • We propose a novel mistake detection framework for assembly tasks. • The framework can better detect ordering mistakes and can be integrated with perception modules. Guodong Ding, Fadime Sener, Shugao Ma, Angela Yao |
Comput. Vis. Image Underst. | 3 |
| 2024 | STMG: A Machine Learning Microgesture Recognition System for Supporting Thumb-Based VR/AR InputabstractAR/VR devices have started to adopt hand tracking, in lieu of controllers, to support user interaction. However, today’s hand input rely primarily on one gesture: pinch. Moreover, current mappings of hand motion to use cases like VR locomotion and content scrolling involve more complex and larger arm motions than joystick or trackpad usage. STMG increases the gesture space by recognizing additional small thumb-based microgestures from skeletal tracking running on a headset. We take a machine learning approach and achieve a 95.1% recognition accuracy across seven thumb gestures performed on the index finger surface: four directional thumb swipes (left, right, forward, backward), thumb tap, and fingertip pinch start and pinch end. We detail the components to our machine learning pipeline and highlight our design decisions and lessons learned in producing a well generalized model. We then demonstrate how these microgestures simplify and reduce arm motions for hand-based locomotion and scrolling interactions. Kenrick Kin, Chengde Wan, Ken Koh, Andrei Marin, Necati Cihan Camgöz, Yujun Cai, Fedor Kovalev, Moshe Ben-Zacharia, Shannon Hoople, Marcos Nunes-Ueno, Mariel Sanchez-Rodriguez, Ayush Bhargava, Robert Wang 0002, Eric Sauser, Shugao Ma |
CHI | 16 |
| 2024 | X-MIC: Cross-Modal Instance Conditioning for Egocentric Action GeneralizationabstractLately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recog-nition. However, the adaptation of these models to egocentric videos has been largely unexplored. To address this gap, we propose a simple yet effective cross-modal adaptation framework, which we call X-MIC. Using a video adapter, our pipeline learns to align frozen text embeddings to each egocentric video directly in the shared embedding space. Our novel adapter architecture retains and improves generalization of the pre-trained VLMs by disentangling learnable temporal modeling and frozen visual en-coder. This results in an enhanced alignment of text embeddings to each egocentric video, leading to a significant improvement in cross-dataset generalization. We evaluate our approach on the Epic-Kitchens, Ego4D, and EGTEA datasets for fine-grained cross-dataset action generalization, demonstrating the effectiveness of our method.11https://github.com/annusha/xmic Anna Kukleva, Fadime Sener, Edoardo Remelli, Bugra Tekin, Eric Sauser, Bernt Schiele, Shugao Ma |
CVPR | 7 |
| 2024 | POET: Prompt Offset Tuning for Continual Human Action Adaptation
Prachi Garg, K. J. Joseph, Vineeth N. Balasubramanian, Necati Cihan Camgöz, Chengde Wan, Kenrick Kin, Weiguang Si, Shugao Ma, Fernando De la Torre |
ECCV (64) | 8 |
| 2024 | On the Utility of 3D Hand Poses for Action Recognition
Md. Salman Shamil, Dibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela Yao |
ECCV (6) | 4 |
| 2024 | DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
Sammy Joe Christen, Shreyas Hampali, Fadime Sener, Edoardo Remelli, Tomas Hodan, Eric Sauser, Shugao Ma, Bugra Tekin |
SIGGRAPH Asia | 7 |
| 2024 | StegoType: Surface Typing from Egocentric CamerasabstractText input is a critical component of any general purpose computing system, yet efficient and natural text input remains a challenge in AR and VR. Headset based hand-tracking has recently become pervasive among consumer VR devices and affords the opportunity to enable touch typing on virtual keyboards. We present an approach for decoding touch typing on uninstrumented flat surfaces using only egocentric camera-based hand-tracking as input. While egocentric hand-tracking accuracy is limited by issues like self occlusion and image fidelity, we show that a sufficiently diverse training set of hand motions paired with typed text can enable a deep learning model to extract signal from this noisy input. Furthermore, by carefully designing a closed-loop data collection process, we can train an end-to-end text decoder that accounts for natural sloppy typing on virtual keyboards. We evaluate our work with a user study (n=18) showing a mean online throughput of 42.4 WPM with an uncorrected error rate (UER) of 7% with our method compared to a physical keyboard baseline of 74.5 WPM at 0.8% UER, showing progress towards unlocking productivity and high throughput use cases in AR/VR. Fadi Botros, Yangyang Shi, Pinhao Guo, Bradford J. Snow, Linguang Zhang, Jingming Dong, Keith Vertanen, Shugao Ma, Robert Wang 0002 |
UIST | 9 |
| 2024 | TouchInsight: Uncertainty-aware Rapid Touch and Text Input for Mixed Reality from Egocentric VisionabstractWhile passive surfaces offer numerous benefits for interaction in mixed reality, reliably detecting touch input solely from head-mounted cameras has been a long-standing challenge. Camera specifics, hand self-occlusion, and rapid movements of both head and fingers introduce considerable uncertainty about the exact location of touch events. Existing methods have thus not been capable of achieving the performance needed for robust interaction. Paul Streli, Fadi Botros, Shugao Ma, Robert Wang 0002, Christian Holz 0001 |
UIST | 4 |
| 2023 | Data-Free Class-Incremental Hand Gesture RecognitionabstractThis paper investigates data-free class-incremental learning (DFCIL) for hand gesture recognition from 3D skeleton sequences. In this class-incremental learning (CIL) setting, while incrementally registering the new classes, we do not have access to the training samples (i.e. data-free) of the already known classes due to privacy. Existing DFCIL methods primarily focus on various forms of knowledge distillation for model inversion to mitigate catastrophic forgetting. Unlike SOTA methods, we delve deeper into the choice of the best samples for inversion. Inspired by the well-grounded theory of max-margin classification, we find that the best samples tend to lie close to the approximate decision boundary within a reasonable margin. To this end, we propose BOAT-MI – a simple and effective boundary-aware prototypical sampling mechanism for model inversion for DFCIL. Our sampling scheme outperforms SOTA methods significantly on two 3D skeleton gesture datasets, the publicly available SHREC 2017, and EgoGesture3D – which we extract from a publicly available RGBD dataset. Both our codebase and the EgoGesture3D skeleton dataset are publicly available: https://github.com/humansensinglab/dfcil-hgr. Shubhra Aich, Jesús Ruiz-Santaquiteria, Prachi Garg, K. J. Joseph, Alvaro Fernandez Garcia, Vineeth N. Balasubramanian, Kenrick Kin, Chengde Wan, Necati Cihan Camgöz, Shugao Ma, Fernando De la Torre |
ICCV | 11 |
| 2023 | Opening the Vocabulary of Egocentric ActionsabstractHuman actions in egocentric videos often feature hand-object interactions composed of a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations — sparsity of action compositions and a closed set of interacting objects. This paper proposes a novel open vocabulary action recognition task. Given a set of verbs and objects observed during training, the goal is to generalize the verbs to an open vocabulary of actions with seen and novel objects. To this end, we decouple the verb and object predictions via an object-agnostic _verb encoder_ and a prompt-based _object encoder_. The prompting leverages CLIP representations to predict an open vocabulary of interacting objects. We create open vocabulary benchmarks on the EPIC-KITCHENS-100 and Assembly101 datasets; whereas closed-action methods fail to generalize, our proposed method is effective. In addition, our object encoder significantly outperforms existing open-vocabulary visual recognition methods in recognizing novel interacting objects. Dibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela Yao |
NeurIPS | 3 |
| 2022 | LiP-Flow: Learning Inference-Time Priors for Codec Avatars via Normalizing Flows in Latent Space
Emre Aksan, Shugao Ma, Akin Caliskan, Stanislav Pidhorskyi, Alexander Richard, Shih-En Wei, Jason M. Saragih, Otmar Hilliges |
ECCV (26) | 2 |
| 2022 | Depth of Field Aware Differentiable RenderingabstractCameras with a finite aperture diameter exhibit defocus for scene elements that are not at the focus distance, and have only a limited depth of field within which objects appear acceptably sharp. In this work we address the problem of applying inverse rendering techniques to input data that exhibits such defocus blurring. We present differentiable depth-of-field rendering techniques that are applicable to both rasterization-based methods using mesh representations, as well as ray-marching-based methods using either explicit [Yu et al. 2021] or implicit volumetric radiance fields [Mildenhall et al. 2020]. Our approach learns significantly sharper scene reconstructions on data containing blur due to depth of field, and recovers aperture and focus distance parameters that result in plausible forward-rendered images. We show applications to macro photography, where typical lens configurations result in a very narrow depth of field, and to multi-camera video capture, where maintaining sharp focus across a large capture volume for a moving subject is difficult. Stanislav Pidhorskyi, Timur M. Bagautdinov, Shugao Ma, Jason M. Saragih, Gabriel Schwartz, Yaser Sheikh, Tomas Simon |
ACM Trans. Graph. | 3 |
| 2021 | Pixel Codec Avatars
Shugao Ma, Tomas Simon, Jason M. Saragih, Yuecheng Li, Fernando De la Torre, Yaser Sheikh |
CVPR | 1 |
| 2021 | F-CAD: A Framework to Explore Hardware Accelerators for Codec Avatar DecodingabstractCreating virtual avatars with realistic rendering is one of the most essential and challenging tasks to provide highly immersive virtual reality (VR) experiences. It requires not only sophisticated deep neural network (DNN) based codec avatar decoders to ensure high visual quality and precise motion expression, but also efficient hardware accelerators to guarantee smooth real-time rendering using lightweight edge devices, like untethered VR headsets. Existing hardware accelerators, however, fail to deliver sufficient performance and efficiency targeting such decoders which consist of multi-branch DNNs and require demanding compute and memory resources. To address these problems, we propose an automation framework, called F-CAD (Facebook Codec avatar Accelerator Design), to explore and deliver optimized hardware accelerators for codec avatar decoding. Novel technologies include 1) a new accelerator architecture to efficiently handle multi-branch DNNs; 2) a multi-branch dynamic design space to enable fine-grained architecture configurations; and 3) an efficient architecture search for picking the optimized hardware design based on both application-specific demands and hardware resource constraints. To the best of our knowledge, F-CAD is the first automation tool that supports the whole design flow of hardware acceleration of codec avatar decoders, allowing joint optimization on decoder designs in popular machine learning frameworks and corresponding customized accelerator design with cycle-accurate evaluation. Results show that the accelerators generated by F-CAD can deliver up to 122.1 frames per second (FPS) and 91.6% hardware efficiency when running the latest codec avatar decoder. Compared to the state-of-the-art designs, F-CAD achieves 4.0× and 2.8× higher throughput, 62.5% and 21.2% higher efficiency than DNNBuilder [1] and HybridDNN [2] by targeting the same hardware device. Xiaofan Zhang 0001, Pierce Chuang, Shugao Ma, Deming Chen, Yuecheng Li |
DAC | 4 |
| 2021 | Audio- and Gaze-driven Facial Animation of Codec AvatarsabstractCodec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are almost indistinguishable from video [28]. In this paper we describe the first approach to animate these parametric models in real-time which could be deployed on commodity virtual reality hardware using audio and/or eye tracking. Our goal is to display expressive conversations between individuals that exhibit important social signals such as laughter and excitement solely from la-tent cues in our lossy input signals. To this end we collected over 5 hours of high frame rate 3D face scans across three participants including traditional neutral speech as well as expressive and conversational speech. We investigate a multimodal fusion approach that dynamically identifies which sensor encoding should animate which parts of the face at any time. See the supplemental video which demonstrates our ability to generate full face motion far beyond the typically neutral lip articulations seen in competing work: https://research.fb.com/videos/audio-and-gaze-driven-facial-animation-of-codec-avatars/. Alexander Richard, Colin Lea, Shugao Ma, Juergen Gall, Fernando De la Torre, Yaser Sheikh |
WACV | 3 |
| 2020 | Expressive Telepresence via Modular Codec Avatars
Hang Chu, Shugao Ma, Fernando De la Torre, Sanja Fidler, Yaser Sheikh |
ECCV (12) | 2 |
| 2019 | Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and SynthesisabstractWe present a 16.2-million frame (50-hour) multimodal dataset of two-person face-to-face spontaneous conversations. Our dataset features synchronized body and finger motion as well as audio data. To the best of our knowledge, it represents the largest motion capture and audio dataset of natural conversations to date. The statistical analysis verifies strong intraperson and interperson covariance of arm, hand, and speech features, potentially enabling new directions on data-driven social behavior analysis, prediction, and synthesis. As an illustration, we propose a novel real-time finger motion synthesis method: a temporal neural network innovatively trained with an inverse kinematics (IK) loss, which adds skeletal structural information to the generative model. Our qualitative user study shows that the finger motion generated by our method is perceived as natural and conversation enhancing, while the quantitative ablation study demonstrates the effectiveness of IK loss. Gilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori, Siddhartha S. Srinivasa, Yaser Sheikh |
ICCV | 3 |
| 2019 | To React or not to React: End-to-End Visual Pose Forecasting for Personalized Avatar during Dyadic ConversationsabstractNon verbal behaviours such as gestures, facial expressions, body posture, and para-linguistic cues have been shown to complement or clarify verbal messages. Hence to improve telepresence, in form of an avatar, it is important to model these behaviours, especially in dyadic interactions. Creating such personalized avatars not only requires to model intrapersonal dynamics between a avatar’s speech and their body pose, but it also needs to model interpersonal dynamics with the interlocutor present in the conversation. In this paper, we introduce a neural architecture named Dyadic Residual-Attention Model (DRAM), which integrates intrapersonal (monadic) and interpersonal (dyadic) dynamics using selective attention to generate sequences of body pose conditioned on audio and body pose of the interlocutor and audio of the human operating the avatar. We evaluate our proposed model on dyadic conversational data consisting of pose and audio of both participants, confirming the importance of adaptive attention between monadic and dyadic dynamics when predicting avatar pose. We also conduct a user study to analyze judgments of human observers. Our results confirm that the generated body pose is more natural, models intrapersonal dynamics and interpersonal dynamics better than non-adaptive monadic/dyadic models. Chaitanya Ahuja, Shugao Ma, Louis-Philippe Morency, Yaser Sheikh |
ICMI | 2 |
| 2018 | Recycle-GAN: Unsupervised Video Retargeting
Aayush Bansal, Shugao Ma, Deva Ramanan, Yaser Sheikh |
ECCV (5) | 2 |
| 2018 | Space-Time Tree Ensemble for Action Recognition and Localization
Shugao Ma, Jianming Zhang 0001, Stan Sclaroff, Nazli Ikizler-Cinbis, Leonid Sigal |
Int. J. Comput. Vis. | 1 |
| 2017 | Salient Object Subitizing
Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech |
Int. J. Comput. Vis. | 2 |
| 2017 | Do less and achieve more: Training CNNs for action recognition utilizing action images from the Web
Shugao Ma, Sarah Adel Bargal, Jianming Zhang 0001, Leonid Sigal, Stan Sclaroff |
Pattern Recognit. | 1 |
| 2016 | Learning Activity Progression in LSTMs for Activity Detection and Early DetectionabstractIn this work we improve training of temporal deep models to better learn activity progression for activity detection and early detection tasks. Conventionally, when training a Recurrent Neural Network, specifically a Long Short Term Memory (LSTM) model, the training loss only considers classification error. However, we argue that the detection score of the correct activity category, or the detection score margin between the correct and incorrect categories, should be monotonically non-decreasing as the model observes more of the activity. We design novel ranking losses that directly penalize the model on violation of such monotonicities, which are used together with classification loss in training of LSTM models. Evaluation on ActivityNet shows significant benefits of the proposed ranking losses in both activity detection and early detection tasks. Shugao Ma, Leonid Sigal, Stan Sclaroff |
CVPR | 1 |
| 2015 | Space-time tree ensemble for action recognitionabstractHuman actions are, inherently, structured patterns of body movements. We explore ensembles of hierarchical spatio-temporal trees, discovered directly from training data, to model these structures for action recognition. The hierarchical spatio-temporal trees provide a robust mid-level representation for actions. However, discovery of frequent and discriminative tree structures is challenging due to the exponential search space, particularly if one allows partial matching. We address this by first building a concise action vocabulary via discriminative clustering. Using the action vocabulary we then utilize tree mining with subsequent tree clustering and ranking to select a compact set of highly discriminative tree patterns. We show that these tree patterns, alone, or in combination with shorter patterns (action words and pairwise patterns) achieve state-of-the-art performance on two challenging datasets: UCF Sports and HighFive. Moreover, trees learned on HighFive are used in recognizing two action classes in a different dataset, Hollywood3D, demonstrating the potential for cross-dataset generality of the trees our approach discovers. Shugao Ma, Leonid Sigal, Stan Sclaroff |
CVPR | 1 |
| 2015 | Salient Object SubitizingabstractPeople can immediately and precisely identify that an image contains 1, 2, 3 or 4 items by a simple glance. The phenomenon, known as Subitizing, inspires us to pursue the task of Salient Object Subitizing (SOS), i.e. predicting the existence and the number of salient objects in a scene using holistic cues. To study this problem, we propose a new image dataset annotated using an online crowdsourcing marketplace. We show that a proposed subitizing technique using an end-to-end Convolutional Neural Network (CNN) model achieves significantly better than chance performance in matching human labels on our dataset. It attains 94% accuracy in detecting the existence of salient objects, and 42–82% accuracy (chance is 20%) in predicting the number of salient objects (1, 2, 3, and 4+), without resorting to any object localization process. Finally, we demonstrate the usefulness of the proposed subitizing technique in two computer vision applications: salient object detection and object proposal. Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech |
CVPR | 2 |
| 2014 | Adaptive Structured Pooling for Action Recognition
Svebor Karaman, Lorenzo Seidenari, Shugao Ma, Alberto Del Bimbo, Stan Sclaroff |
BMVC | 3 |
| 2014 | MEEM: Robust Tracking via Multiple Experts Using Entropy Minimization
Jianming Zhang 0001, Shugao Ma, Stan Sclaroff |
ECCV (6) | 2 |
| 2013 | Action Recognition and Localization by Hierarchical Space-Time SegmentsabstractWe propose Hierarchical Space-Time Segments as a new representation for action recognition and localization. This representation has a two-level hierarchy. The first level comprises the root space-time segments that may contain a human body. The second level comprises multi-grained space-time segments that contain parts of the root. We present an unsupervised method to generate this representation from video, which extracts both static and non-static relevant space-time segments, and also preserves their hierarchical and temporal relationships. Using simple linear SVM on the resultant bag of hierarchical space-time segments representation, we attain better than, or comparable to, state-of-the-art action recognition performance on two challenging benchmark datasets and at the same time produce good action localization results. Shugao Ma, Jianming Zhang 0001, Nazli Ikizler-Cinbis, Stan Sclaroff |
ICCV | 1 |
| 2008 | Effective scene matching with local feature representativesabstractScene matching measures the similarity of scenes in photos and is of central importance in applications where we have to properly organize large amount of digital photos by scene categories. In this paper, we present a novel scene matching method using local features representatives. For a given image, its scene is compactly represented as a set of cluster centers, called local feature representatives, where the clusters are obtained using the affinity propagation (AP) algorithm to aggregate local features according to their spatial closeness and appearance similarity. The similarity of scenes in two images is then measured by a modified Earth Mover Distance (EMD) between their corresponding sets of local feature representatives. Empirical experiments on real world photos shows that our method is comparable to the state-of-the-arts. Shugao Ma, Weiqiang Wang 0001, Qingming Huang, Shuqiang Jiang, Wen Gao 0001 |
ICPR | 1 |