EDBT 2026 Demo / reviewers in the wild / expert
Michael S. Ryoo
dblp:r/MichaelSRyoo
· DBLP profile ↗
95ranked-venue papers
24as first author
38since 2021 · last 2025
0000-0002-5452-8332ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 82 · 21 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 57 · 15 first-author · 22 since 2021Systems, architecture and hardware · 12 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped NoiseabstractGenerative modeling aims to transform random noise into structured outputs. In this work, we enhance video diffusion models by allowing motion control via structured latent noise sampling. This is achieved by just a change in data: we pre-process training videos to yield structured noise. Consequently, our method is agnostic to diffusion model design, requiring no changes to model architectures or training pipelines. Specifically, we propose a novel noise warping algorithm, fast enough to run in real time, that replaces random temporal Gaussianity with correlated warped noise derived from optical flow fields, while preserving the spatial Gaussianity. The efficiency of our algorithm enables us to fine-tune modern video diffusion base models using warped noise with minimal overhead, and provide a one-stop solution for a wide range of userfriendly motion control: local object motion control, global camera movement control, and motion transfer. The harmonization between temporal coherence and spatial Gaussianity in our warped noise leads to effective motion control while maintaining per-frame pixel quality. Extensive experiments and user studies demonstrate the advantages of our method, making it a robust and scalable approach for controlling motion in video diffusion models. Please see our project webpage; source code and checkpoints are available on GitHub. Ryan D. Burgert, Yuancheng Xu, Wenqi Xian, Oliver Pilarski, Pascal Clausen, Mingming He, Yitong Deng, Mohsen Mousavi, Michael S. Ryoo, Paul E. Debevec, Ning Yu 0006 |
CVPR | 11 |
| 2025 | Adaptive Caching for Faster Video Generation With Diffusion TransformersabstractGenerating temporally-consistent high-fidelity videos can be computationally expensive, especially over longer temporal spans. More-recent Diffusion Transformers (DiTs) -- despite making significant headway in this context -- have only heightened such challenges as they rely on larger models and heavier attention mechanisms, resulting in slower inference speeds. In this paper, we introduce a training-free method to accelerate video DiTs, termed Adaptive Caching (AdaCache), which is motivated by the fact that "not all videos are created equal": meaning, some videos require fewer denoising steps to attain a reasonable quality than others. Building on this, we not only cache computations through the diffusion process, but also devise a caching schedule tailored to each video generation, maximizing the quality-latency trade-off. We further introduce a Motion Regularization (MoReg) scheme to utilize video information within AdaCache, essentially controlling the compute allocation based on motion content. Altogether, our plug-and-play contributions grant significant inference speedups (e.g. up to 4.7x on Open-Sora 720p - 2s video generation) without sacrificing the generation quality, across multiple video DiT baselines. Kumara Kahatapitiya, Sen He 0001, Menglin Jia, Michael S. Ryoo, Tian Xie 0003 |
ICCV | 7 |
| 2025 | LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyabstractVision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for robotic control remains challenging, particularly when constrained by a limited number of robot demonstrations. In this work, we introduce LLaRA: Large Language and Robotics Assistant, a framework that formulates robot action policy as visuo-textual conversations and enables an efficient transfer of a pretrained VLM into a powerful VLA, motivated by the success of visual instruction tuning in Computer Vision. First, we present an automated pipeline to generate conversation-style instruction tuning data for robots from existing behavior cloning datasets, aligning robotic actions with image pixel coordinates. Further, we enhance this dataset in a self-supervised manner by defining six auxiliary tasks, without requiring any additional action annotations. We show that a VLM finetuned with a limited amount of such datasets can produce meaningful action decisions for robotic control. Through experiments across multiple simulated and real-world tasks, we demonstrate that LLaRA achieves state-of-the-art performance while preserving the generalization capabilities of large language models. The code, datasets, and pretrained models are available at https://github.com/LostXine/LLaRA. Xiang Li 0109, Cristina Mata, Jongwoo Park 0003, Kumara Kahatapitiya, Yoo Sung Jang, Jinghuan Shang, Kanchana Ranasinghe, Ryan D. Burgert, Mu Cai, Yong Jae Lee, Michael S. Ryoo |
ICLR | 11 |
| 2025 | Understanding Long Videos with Multimodal Language ModelsabstractLarge Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs influence this strong performance. Surprisingly, we discover that LLM-based approaches can yield surprisingly good accuracy on long-video tasks with limited video information, sometimes even with no video-specific information. Building on this, we explore injecting video-specific information into an LLM-based framework. We utilize off-the-shelf vision tools to extract three object-centric information modalities from videos, and then leverage natural language as a medium for fusing this information. Our resulting Multimodal Video Understanding (MVU) framework demonstrates state-of-the-art performance across multiple video understanding benchmarks. Strong performance also on robotics domain tasks establishes its strong generality. Code: github.com/kahnchana/mvu Kanchana Ranasinghe, Xiang Li 0109, Kumara Kahatapitiya, Michael S. Ryoo |
ICLR | 4 |
| 2024 | MAGICK: A Large-Scale Captioned Dataset from Matting Generated Images Using Chroma KeyingabstractWe introduce MAGICK, a large-scale dataset of generated objects with high-quality alpha mattes. While image generation methods have produced segmentations, they cannot generate alpha mattes with accurate details in hair, fur, and transparencies. This is likely due to the small size of current alpha matting datasets and the difficulty in obtaining ground-truth alpha. We propose a scalable method for synthesizing images of objects with high-quality alpha that can be used as a ground-truth dataset. A key idea is to generate objects on a single-colored background so chroma keying approaches can be used to extract the alpha. However, this faces several challenges, including that current text-to-image generation methods cannot create images that can be easily chroma keyed and that chroma keying is an underconstrained problem that generally requires manual intervention for high-quality results. We address this using a combination of generation and alpha extraction methods. Using our method, we generate a dataset of 150,000 objects with alpha. We show the utility of our dataset by training an alpha-to-rgb generation method that outperforms baselines. Please see our project website at https://ryanndagreat.github.io/MAGICK/. Ryan D. Burgert, Brian L. Price, Jason Kuen, Michael S. Ryoo |
CVPR | 5 |
| 2024 | VicTR: Video-conditioned Text Representations for Activity RecognitionabstractVision-Language models (VLMs) have excelled in the image-domain- especially in zero-shot settings- thanks to the availability of vast pretraining data (i.e., paired image-text samples). However for videos, such paired data is not as abundant. Therefore, video- VLMs are usually designed by adapting pretrained image- VLMs to the video-domain, instead of training from scratch. All such recipes rely on aug-menting visual embeddings with temporal information (i.e., image -+ video), often keeping text embeddings unchanged or even being discarded. In this paper, we argue the contrary, that better video- VLMs can be designed by focusing more on augmenting text, rather than visual information. More specifically, we introduce Video-conditioned Text Representations (Vi c TR): a form of text embeddings optimized w.r.t. vi-sual embeddings, creating a more-flexible contrastive latent space. Our model canfurther make use offreely-available semantic information, in the form of visually- grounded aux-iliary text (e.g. object or scene information). We evaluate our model on few-shot, zero-shot (HMDB-51, UCF-10l), short-form (Kinetics-400) and long-form (Charades) activ-ity recognition benchmarks, showing strong performance among video-VLMs. Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. Ryoo |
CVPR | 4 |
| 2024 | Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual ModalitiesabstractOne of the main challenges of multimodal learning is combining multiple heterogeneous modalities, e.g., video, audio, and text. Video and audio are obtained at much higher rates than text and are roughly aligned in time. They are often not synchronized with text, which comes as a global context, e.g. a title, or a description. Furthermore, video and audio inputs are of much larger volumes, and grow as the video length increases, which naturally requires more compute dedicated to these modalities, and makes modeling of long-range dependencies harder. We here decouple the multimodal modeling, dividing it into separate autoregressive models, processing the inputs according to the characteristics of the modalities. We propose a multimodal model, consisting of an autoregressive component for the time-synchronized modalities (audio and video), and an autoregressive component for the context modalities which are not necessarily aligned in time but are still sequential. To address the long-sequences of the video-audio inputs, we further partition the video and audio sequences in consecutive snippets and autoregressively process their representations. To that end, we propose a Combiner mechanism, which models the audio-video information jointly, producing compact but expressive representations. This allows us to scale to 512 input video frames without increase in model parameters. Our approach achieves the state-of-the-art on multiple well established multimodal benchmarks. It effectively addresses the high computational demand of media inputs by learning compact representations, controlling the sequence length of the audio-video feature representations, and modeling their dependencies in time. A. J. Piergiovanni, Isaac Noble, Dahun Kim, Michael S. Ryoo, Victor Gomes, Anelia Angelova |
CVPR | 4 |
| 2024 | Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMsabstractIntegration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly for visual question answering (VQA). However, existing V-LLMs (e.g. BLIP-2, LLaVA) demonstrate weak spatial reasoning and localization awareness. Despite generating highly descriptive and elaborate textual answers, these models fail at simple tasks like distinguishing a left vs right location. In this work, we explore how image-space coordinate based instruction fine-tuning objectives could inject spatial awareness into V-LLMs. We discover optimal coordinate representations, data-efficient instruction fine-tuning objectives, and pseudo-data generation strategies that lead to improved spatial awareness in V-LLMs. Additionally, our resulting model improves VQA across image and video domains, reduces undesired hallucination, and generates better contextual object descriptions. Experiments across 5 vision-language tasks involving 14 different datasets establish the clear performance improvements achieved by our proposed framework. Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo, Tsung-Yu Lin |
CVPR | 4 |
| 2024 | CoPT: Unsupervised Domain Adaptive Segmentation Using Domain-Agnostic Text Embeddings
Cristina Mata, Kanchana Ranasinghe, Michael S. Ryoo |
ECCV (62) | 3 |
| 2024 | SARA-RT: Scaling up Robotics Transformers with Self-Adaptive Robust AttentionabstractWe present Self-Adaptive Robust Attention for Robotics Transformers (SARA-RT): a new paradigm for addressing the emerging challenge of scaling up Robotics Transformers (RT) for on-robot deployment. SARA-RT relies on the new method of fine-tuning proposed by us, called up-training. It converts pre-trained or already fine-tuned Transformer-based robotic policies of quadratic time complexity (including massive billion-parameter vision-language-action models or VLAs), into their efficient linear-attention counterparts maintaining high quality. We demonstrate the effectiveness of SARA-RT by speeding up: (a) the class of recently introduced RT-2 models [1], the first VLA robotic policies pre-trained on internet-scale data, as well as (b) Point Cloud Transformer (PCT) robotic policies operating on large point clouds. We complement our results with the rigorous mathematical analysis providing deeper insight into the phenomenon of SARA. Isabel Leal, Krzysztof Choromanski, Deepali Jain, Avinava Dubey, Jake Varley, Michael S. Ryoo, Yao Lu 0006, Frederick Liu, Vikas Sindhwani, Tamás Sarlós, Kenneth Oslund, Karol Hausman, Kanishka Rao |
ICRA | 6 |
| 2024 | Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised LearningabstractDiffusion models have been adopted for behavioral cloning in a sequence modeling fashion, benefiting from their exceptional capabilities in modeling complex data distributions. The standard diffusion-based policy iteratively denoises action sequences from random noise conditioned on the input states and the model is typically trained with a singular diffusion loss. This paper explores the potential enhancements in such models when the denoising process is informed by a better visual representation. We study the scenario where the model is jointly optimized using the standard diffusion loss alongside an auxiliary objective based on self-supervised learning. After experimenting with various objectives, we introduce Crossway Diffusion, a simple yet effective way to enhance diffusion-based visuomotor policy learning via a state decoder and an auxiliary reconstruction objective. During training, the state decoder reconstructs raw image pixels and other states from the intermediate representations of the model. Experiments demonstrate the effectiveness of our method in various simulated and real-world tasks, confirming its consistent advantages over the standard diffusion-based policy and other baselines. Xiang Li 0109, Varun Belagali, Jinghuan Shang, Michael S. Ryoo |
ICRA | 4 |
| 2024 | Grafting Vision TransformersabstractVision Transformers (ViTs) have recently become the state-of-the-art across many computer vision tasks. In contrast to convolutional networks (CNNs), ViTs enable global information sharing even within shallow layers of a network, i.e., among high-resolution features. However, this perk was later overlooked with the success of pyramid architectures such as Swin Transformer, which show better performance-complexity trade-offs. In this paper, we present a simple and efficient add-on component (termed GrafT) that considers global dependencies and multi-scale information throughout the network, in both high- and low-resolution features alike. It has the flexibility of branching out at arbitrary depths and shares most of the parameters and computations of the backbone. GrafT shows consistent gains over various well-known models which includes both hybrid and pure Transformer types, both homogeneous and pyramid structures, and various self-attention methods. In particular, it largely benefits mobile-size models by providing high-level semantics. On the ImageNet-1k dataset, GrafT delivers +3.9%, +1.4%, and +1.9% top-1 accuracy improvement to DeiT-T, Swin-T, and MobViTXXS, respectively. Our code and models are at https://github.com/jongwoopark7978/Grafting-Vision-Transformer. Jongwoo Park 0003, Kumara Kahatapitiya, Donghyun Kim 0006, Shivchander Sudalairaj, Quanfu Fan, Michael S. Ryoo |
WACV | 6 |
| 2024 | Limited Data, Unlimited Potential: A Study on ViTs Augmented by Masked AutoencodersabstractVision Transformers (ViTs) have become ubiquitous in computer vision. Despite their success, ViTs lack inductive biases, which can make it difficult to train them with limited data. To address this challenge, prior studies suggest training ViTs with self-supervised learning (SSL) and fine-tuning sequentially. However, we observe that jointly optimizing ViTs for the primary task and a Self-Supervised Auxiliary Task (SSAT) is surprisingly beneficial when the amount of training data is limited. We explore the appropriate SSL tasks that can be optimized alongside the primary task, the training schemes for these tasks, and the data scale at which they can be most effective. Our findings reveal that SSAT is a powerful technique that enables ViTs to leverage the unique characteristics of both the self-supervised and primary tasks, achieving better performance than typical ViTs pre-training with SSL and fine-tuning sequentially. Our experiments, conducted on 10 datasets, demonstrate that SSAT significantly improves ViT performance while reducing carbon footprint. We also confirm the effectiveness of SSAT in the video domain for deepfake detection, showcasing its generalizability. Our code is available at https://github.com/dominickrei/Limited-data-vits. Srijan Das, Tanmay Jain, Dominick Reilly, Pranav Balaji, Soumyajit Karmakar, Shyam Marjit, Xiang Li 0109, Abhijit Das 0001, Michael S. Ryoo |
WACV | 9 |
| 2023 | Weakly-Guided Self-Supervised Pretraining for Temporal Activity DetectionabstractTemporal Activity Detection aims to predict activity classes per frame, in contrast to video-level predictions in Activity Classification (i.e., Activity Recognition). Due to the expensive frame-level annotations required for detection, the scale of detection datasets is limited. Thus, commonly, previous work on temporal activity detection resorts to fine-tuning a classification model pretrained on large-scale classification datasets (e.g., Kinetics-400). However, such pretrained models are not ideal for downstream detection, due to the disparity between the pretraining and the downstream fine-tuning tasks. In this work, we propose a novel weakly-guided self-supervised pretraining method for detection. We leverage weak labels (classification) to introduce a self-supervised pretext task (detection) by generating frame-level pseudo labels, multi-action frames, and action segments. Simply put, we design a detection task similar to downstream, on large-scale classification data, without extra annotations. We show that the models pretrained with the proposed weakly-guided self-supervised detection task outperform prior work on multiple challenging activity detection benchmarks, including Charades and MultiTHUMOS. Our extensive ablations further provide insights on when and how to use the proposed models for activity detection. Code is available at github.com/kkahatapitiya/SSDet. Kumara Kahatapitiya, Zhou Ren, Zhenyu Wu 0002, Michael S. Ryoo, Gang Hua 0001 |
AAAI | 5 |
| 2023 | Attributes-Aware Network for Temporal Action Detection
Rui Dai 0001, Srijan Das, Michael S. Ryoo, François Brémond |
BMVC | 3 |
| 2023 | Token Turing MachinesabstractWe propose Token Turing Machines (TTM), a sequential, autoregressive Transformer model with memory for real-world sequential visual understanding. Our model is inspired by the seminal Neural Turing Machine, and has an external memory consisting of a set of tokens which summarise the previous history (i.e., frames). This memory is efficiently addressed, read and written using a Transformer as the processing unit/controller at each step. The model's memory module ensures that a new observation will only be processed with the contents of the memory (and not the entire history), meaning that it can efficiently process long sequences with a bounded computational cost at each step. We show that TTM outperforms other alternatives, such as other Transformer models designed for long sequences and recurrent neural networks, on two real-world sequential visual understanding tasks: online temporal activity detection from videos and vision-based robot action policy learning. Code is publicly available at: https://github.com/google-research/scenic/tree/main/scenic/projects/token.turing. Michael S. Ryoo, Keerthana Gopalakrishnan, Kumara Kahatapitiya, Ted Xiao, Kanishka Rao, Austin Stone, Yao Lu 0006, Julian Ibarz, Anurag Arnab |
CVPR | 1 |
| 2023 | Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Andy Zeng 0001, Maria Attarian, Brian Ichter, Krzysztof Choromanski, Adrian Wong, Stefan Welker, Federico Tombari, Aveek Purohit, Michael S. Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, Peter R. Florence |
ICLR | 9 |
| 2023 | Open-vocabulary Queryable Scene Representations for Real World PlanningabstractLarge language models (LLMs) have unlocked new capabilities of task planning from human instructions. However, prior attempts to apply LLMs to real-world robotic tasks are limited by the lack of grounding in the surrounding scene. In this paper, we develop NLMap, an open-vocabulary and queryable scene representation to address this problem. NLMap serves as a framework to gather and integrate contextual information into LLM planners, allowing them to see and query available objects in the scene before generating a context-conditioned plan. NLMap first establishes a natural language queryable scene representation with Visual Language models (VLMs). An LLM based object proposal module parses instructions and proposes involved objects to query the scene representation for object availability and location. An LLM planner then plans with such information about the scene. NLMap allows robots to operate without a fixed list of objects nor executable options, enabling real robot operation unachievable by previous methods. Project website: https://nlmap-saycan.github.io. Boyuan Chen 0003, Fei Xia 0002, Brian Ichter, Kanishka Rao, Keerthana Gopalakrishnan, Michael S. Ryoo, Austin Stone, Daniel Kappler |
ICRA | 6 |
| 2023 | Energy-Based Models for Cross-Modal Localization using Convolutional TransformersabstractWe present a novel framework using Energy-Based Models (EBMs) for localizing a ground vehicle mounted with a range sensor against satellite imagery in the absence of GPS. Lidar sensors have become ubiquitous on autonomous vehicles for describing its surrounding environment. Map priors are typically built using the same sensor modality for localization purposes. However, these map building endeavors using range sensors are often expensive and time-consuming. Alternatively, we leverage the use of satellite images as map priors, which are widely available, easily accessible, and pro-vide comprehensive coverage. We propose a method using convolutional transformers that performs accurate metric-level localization in a cross-modal manner, which is challenging due to the drastic difference in appearance between the sparse range sensor readings and the rich satellite imagery. We train our model end-to-end and demonstrate our approach achieving higher accuracy than the state-of-the-art on KITTI, Pandaset, and a custom dataset. Michael S. Ryoo |
ICRA | 2 |
| 2023 | SWAT: Spatial Structure Within and Among TokensabstractModeling visual data as tokens (i.e., image patches) using attention mechanisms, feed-forward networks or convolutions has been highly effective in recent years. Such methods usually have a common pipeline: a tokenization method, followed by a set of layers/blocks for information mixing, both within and among tokens. When image patches are converted into tokens, they are often flattened, discarding the spatial structure within each patch. As a result, any processing that follows (eg: multi-head self-attention) may fail to recover and/or benefit from such information. In this paper, we argue that models can have significant gains when spatial structure is preserved during tokenization, and is explicitly used during the mixing stage. We propose two key contributions: (1) Structure-aware Tokenization and, (2) Structure-aware Mixing, both of which can be combined with existing models with minimal effort. We introduce a family of models (SWAT), showing improvements over the likes of DeiT, MLP-Mixer and Swin Transformer, across multiple benchmarks including ImageNet classification and ADE20K segmentation. Our code is available at github.com/kkahatapitiya/SWAT. Kumara Kahatapitiya, Michael S. Ryoo |
IJCAI | 2 |
| 2023 | Language-based Action Concept Spaces Improve Video Self-Supervised LearningabstractRecent contrastive language image pre-training has led to learning highly transferable and robust image representations. However, adapting these models to video domain with minimal supervision remains an open problem. We explore a simple step in that direction, using language tied self-supervised learning to adapt an image CLIP model to the video domain. A backbone modified for temporal modeling is trained under self-distillation settings with train objectives operating in an action concept space. Feature vectors of various action concepts extracted from a language encoder using relevant textual prompts construct this space. A large language model aware of actions and their attributes generates the relevant textual prompts.
We introduce two train objectives, concept distillation and concept alignment, that retain generality of original representations while enforcing relations between actions and their attributes. Our approach improves zero-shot and linear probing performance on three action recognition benchmarks. Kanchana Ranasinghe, Michael S. Ryoo |
NeurIPS | 2 |
| 2023 | Active Vision Reinforcement Learning under Limited Visual ObservabilityabstractIn this work, we investigate Active Vision Reinforcement Learning (ActiveVision-RL), where an embodied agent simultaneously learns action policy for the task while also controlling its visual observations in partially observable environments. We denote the former as motor policy and the latter as sensory policy. For example, humans solve real world tasks by hand manipulation (motor policy) together with eye movements (sensory policy). ActiveVision-RL poses challenges on coordinating two policies given their mutual influence. We propose SUGARL, Sensorimotor Understanding Guided Active Reinforcement Learning, a framework that models motor and sensory policies separately, but jointly learns them using with an intrinsic sensorimotor reward. This learnable reward is assigned by sensorimotor reward module, incentivizes the sensory policy to select observations that are optimal to infer its own motor action, inspired by the sensorimotor stage of humans. Through a series of experiments, we show the effectiveness of our method across a range of observability conditions and its adaptability to existed RL algorithms. The sensory policies learned through our method are observed to exhibit effective active vision strategies. Jinghuan Shang, Michael S. Ryoo |
NeurIPS | 2 |
| 2023 | ViewCLR: Learning Self-supervised Video Representation for Unseen ViewpointsabstractLearning self-supervised video representation predominantly focuses on discriminating instances generated from simple data augmentation schemes. However, the learned representation often fails to generalize over unseen camera viewpoints. To this end, we propose ViewCLR, that learns self-supervised video representation invariant to camera viewpoint changes. We introduce a viewpoint-generator that can be considered as a learnable augmentation for any self-supervised pre-text tasks, to generate latent viewpoint representation of a video. ViewCLR maximizes the similarities between the representation of the latent viewpoint and that of the original viewpoint, enabling the learned video encoder to generalize over unseen camera viewpoints. Experiments on cross-view benchmark datasets including NTU RGB+D dataset show that ViewCLR stands as a state-of-the-art viewpoint invariant self-supervised method. Srijan Das, Michael S. Ryoo |
WACV | 2 |
| 2023 | StARformer: Transformer With State-Action-Reward Representations for Robot LearningabstractReinforcement Learning (RL) can be considered as a sequence modeling task, where an agent employs a sequence of past state-action-reward experiences to predict a sequence of future actions. In this work, we propose State-Action-Reward Transformer (StARformer), a Transformer architecture for robot learning with image inputs, which explicitly models short-term state-action-reward representations (StAR-representations), essentially introducing a Markovian-like inductive bias to improve long-term modeling. StARformer first extracts StAR-representations using self-attending patches of image states, action, and reward tokens within a short temporal window. These StAR-representations are combined with pure image state representations, extracted as convolutional features, to perform self-attention over the whole sequence. Our experimental results show that StARformer outperforms the state-of-the-art Transformer-based method on image-based Atari and DeepMind Control Suite benchmarks, under both offline-RL and imitation learning settings. We find that models can benefit from our combination of patch-wise and convolutional image embeddings. StARformer is also more compliant with longer sequences of inputs than the baseline method. Finally, we demonstrate how StARformer can be successfully applied to a real-world robot imitation learning setting via a human-following task. Jinghuan Shang, Xiang Li 0109, Kumara Kahatapitiya, Yu-Cheol Lee, Michael S. Ryoo |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionabstractAction detection is a significant and challenging task, especially in densely-labelled datasets of untrimmed videos. Such data consist of complex temporal relations including composite or co-occurring actions. To detect actions in these complex settings, it is critical to capture both shortterm and long-term temporal information efficiently. To this end, we propose a novel ‘ConvTransformer’ network for action detection: MS-TCT11Code/Models: https://github.com/dairui01/MS-TCT. This network comprises of three main components: (1) a Temporal Encoder module which explores global and local temporal relations at multiple temporal resolutions, (2) a Temporal Scale Mixer module which effectively fuses multi-scale features, creating a unified feature representation, and (3) a Classification module which learns a center-relative position of each action instance in time, and predicts frame-level classification scores. Our experimental results on multiple challenging datasets such as Charades, TSU and MultiTHUMOS, validate the effectiveness of the proposed method, which outperforms the state-of-the-art methods on all three datasets. Rui Dai 0001, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo, François Brémond |
CVPR | 4 |
| 2022 | Self-supervised Video TransformerabstractIn this paper, we propose self-supervised training for video transformers using unlabeled video data. From a given video, we create local and global spatiotemporal views with varying spatial sizes and frame rates. Our self-supervised objective seeks to match the features of these different views representing the same video, to be invariant to spatiotemporal variations in actions. To the best of our knowledge, the proposed approach is the first to alleviate the dependency on negative samples or dedicated memory banks in Self-supervised Video Transformer (SVT). Further, owing to the flexibility of Transformer models, SVT supports slow-fast video processing within a single architecture using dynamically adjusted positional encoding and supports longterm relationship modeling along spatiotemporal dimensions. Our approach performs well on four action recognition benchmarks (Kinetics-400, UCF-101, HMDB-51, and SSv2) and converges faster with small batch sizes. Code is available at: https://git.io/J1juJ. Kanchana Ranasinghe, Muzammal Naseer, Salman Khan 0001, Fahad Shahbaz Khan, Michael S. Ryoo |
CVPR | 5 |
| 2022 | Video Question Answering with Iterative Video-Text Co-tokenization
A. J. Piergiovanni, Kairo Morton, Weicheng Kuo, Michael S. Ryoo, Anelia Angelova |
ECCV (36) | 4 |
| 2022 | StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning
Jinghuan Shang, Kumara Kahatapitiya, Xiang Li 0109, Michael S. Ryoo |
ECCV (39) | 4 |
| 2022 | Hybrid Random Features
Krzysztof Choromanski, Haoxian Chen 0002, Arijit Sehanobish, Yuanzhe Ma, Deepali Jain, Jake Varley, Andy Zeng 0001, Michael S. Ryoo, Valerii Likhosherstov, Dmitry Kalashnikov, Vikas Sindhwani, Adrian Weller |
ICLR | 9 |
| 2022 | Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?abstractWe investigate whether self-supervised learning (SSL) can improve online reinforcement learning (RL) from pixels. We extend the contrastive reinforcement learning framework (e.g., CURL) that jointly optimizes SSL and RL losses and conduct an extensive amount of experiments with various self-supervised losses. Our observations suggest that the existing SSL framework for RL fails to bring meaningful improvement over the baselines only taking advantage of image augmentation when the same amount of data and augmentation is used. We further perform evolutionary searches to find the optimal combination of multiple self-supervised losses for RL, but find that even such a loss combination fails to meaningfully outperform the methods that only utilize carefully designed image augmentations. After evaluating these approaches together in multiple different environments including a real-world robot environment, we confirm that no single self-supervised loss or image augmentation method can dominate all environments and that the current framework for joint optimization of SSL and RL is limited. Finally, we conduct the ablation study on multiple factors and demonstrate the properties of representations learned with different approaches. Xiang Li 0109, Jinghuan Shang, Srijan Das, Michael S. Ryoo |
NeurIPS | 4 |
| 2022 | Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceabstractHumans are remarkably flexible in understanding viewpoint changes due to visual cortex supporting the perception of 3D structure. In contrast, most of the computer vision models that learn visual representation from a pool of 2D images often fail to generalize over novel camera viewpoints. Recently, the vision architectures have shifted towards convolution-free architectures, visual Transformers, which operate on tokens derived from image patches. However, these Transformers do not perform explicit operations to learn viewpoint-agnostic representation for visual understanding. To this end, we propose a 3D Token Representation Layer (3DTRL) that estimates the 3D positional information of the visual tokens and leverages it for learning viewpoint-agnostic representations. The key elements of 3DTRL include a pseudo-depth estimator and a learned camera matrix to impose geometric transformations on the tokens, trained in an unsupervised fashion. These enable 3DTRL to recover the 3D positional information of the tokens from 2D patches. In practice, 3DTRL is easily plugged-in into a Transformer. Our experiments demonstrate the effectiveness of 3DTRL in many vision tasks including image classification, multi-view video alignment, and action recognition. The models with 3DTRL outperform their backbone Transformers in all the tasks with minimal added computation. Our code is available at https://github.com/elicassion/3DTRL. Jinghuan Shang, Srijan Das, Michael S. Ryoo |
NeurIPS | 3 |
| 2021 | Unsupervised Discovery of Actions in Instructional Videos
A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo, Irfan A. Essa |
BMVC | 3 |
| 2021 | Coarse-Fine Networks for Temporal Activity Detection in VideosabstractIn this paper, we introduce Coarse-Fine Networks, a twostream architecture which benefits from different abstractions of temporal resolution to learn better video representations for long-term motion. Traditional Video models process inputs at one (or few) fixed temporal resolution without any dynamic frame selection. However, we argue that, processing multiple temporal resolutions of the input and doing so dynamically by learning to estimate the importance of each frame can largely improve video representations, specially in the domain of temporal activity localization. To this end, we propose (1) ‘Grid Pool’, a learned temporal downsampling layer to extract coarse features, and, (2) ‘Multi-stage Fusion’, a spatio-temporal attention mechanism to fuse a finegrained context with the coarse features. We show that our method outperforms the state-of-the-arts for action detection in public datasets including Charades with a significantly reduced compute and memory footprint. The code is available at https://github.com/kkahatapitiya/Coarse-Fine-Networks. Kumara Kahatapitiya, Michael S. Ryoo |
CVPR | 2 |
| 2021 | Recognizing Actions in Videos From Unseen ViewpointsabstractStandard methods for video recognition use large CNNs designed to capture spatio-temporal data. However, training these models requires a large amount of labeled training data, containing a wide variety of actions, scenes, settings and camera viewpoints. In this paper, we show that current convolutional neural network models are unable to recognize actions from camera viewpoints not present in their training data (i.e., unseen view action recognition). To address this, we develop approaches based on 3D representations and introduce a new geometric convolutional layer that can learn viewpoint invariant representations. Further, we introduce a new, challenging dataset for unseen view recognition and show the approaches ability to learn viewpoint invariant representations. A. J. Piergiovanni, Michael S. Ryoo |
CVPR | 2 |
| 2021 | 4D-Net for Learned Multi-Modal AlignmentabstractWe present 4D-Net, a 3D object detection approach, which utilizes 3D Point Cloud and RGB sensing information, both in time. We are able to incorporate the 4D information by performing a novel dynamic connection learning across various feature representations and levels of abstraction, as well as by observing geometric constraints. Our approach outperforms the state-of-the-art and strong base-lines on the Waymo Open Dataset. 4D-Net is better able to use motion cues and dense image information to detect distant objects more successfully. We will open source the code. A. J. Piergiovanni, Vincent Casser, Michael S. Ryoo, Anelia Angelova |
ICCV | 3 |
| 2021 | Visionary: Vision architecture discovery for robot learningabstractWe propose a vision-based architecture search algorithm for robot manipulation learning, which discovers interactions between low dimension action inputs and high dimensional visual inputs. Our approach automatically designs architectures while training on the task – discovering novel ways of combining and attending image feature representations with actions as well as features from previous layers. The obtained new architectures demonstrate better task success rates, in some cases with a large margin, compared to a recent high performing baseline. Our real robot experiments also confirm that it improves grasping performance by 6%. This is the first approach to demonstrate a successful neural architecture search and attention connectivity search for a real-robot task. Iretiayo Akinola, Anelia Angelova, Yao Lu 0006, Yevgen Chebotar, Dmitry Kalashnikov, Jacob Varley, Julian Ibarz, Michael S. Ryoo |
ICRA | 8 |
| 2021 | Self-Supervised Disentangled Representation Learning for Third-Person Imitation LearningabstractHumans learn to imitate by observing others. However, robot imitation learning generally requires expert demonstrations in the first-person view (FPV). Collecting such FPV videos for every robot could be very expensive.Third-person imitation learning (TPIL) is the concept of learning action policies by observing other agents in a third-person view (TPV), similar to what humans do. This ultimately allows utilizing human and robot demonstration videos in TPV from many different data sources, for the policy learning. In this paper, we present a TPIL approach for robot tasks with egomotion. Although many robot tasks with ground/aerial mobility often involve actions with camera egomotion, study on TPIL for such tasks has been limited. Here, FPV and TPV observations are visually very different; FPV shows egomotion while the agent appearance is only observable in TPV. To enable better state learning for TPIL, we propose our disentangled representation learning method. We use a dual auto-encoder structure plus representation permutation loss and time-contrastive loss to ensure the state and viewpoint representations are well disentangled. Our experiments show the effectiveness of our approach. Jinghuan Shang, Michael S. Ryoo |
IROS | 2 |
| 2021 | TokenLearner: Adaptive Space-Time Tokenization for VideosabstractIn this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks. Instead of relying on hand-designed splitting strategies to obtain visual tokens and processing a large number of densely sampled patches for attention, our approach learns to mine important tokens in visual data. This results in efficiently and effectively finding a few important visual tokens and enables modeling of pairwise attention between such tokens, over a longer temporal horizon for videos, or the spatial content in image frames. Our experiments demonstrate strong performance on several challenging benchmarks for video recognition tasks. Importantly, due to our tokens being adaptive, we accomplish competitive results at significantly reduced computational cost. We establish new state-of-the-arts on multiple video datasets, including Kinetics-400, Kinetics-600, Charades, and AViD. Michael S. Ryoo, A. J. Piergiovanni, Anurag Arnab, Mostafa Dehghani 0001, Anelia Angelova |
NeurIPS | 1 |
| 2020 | Differentiable Grammars for Videos
A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo |
AAAI | 3 |
| 2020 | Evolving Losses for Unsupervised Video Representation LearningabstractWe present a new method to learn video representations from large-scale unlabeled video data. Ideally, this representation will be generic and transferable, directly usable for new tasks such as action recognition and zero or few-shot learning. We formulate unsupervised representation learning as a multi-modal, multi-task learning problem, where the representations are shared across different modalities via distillation. Further, we introduce the concept of loss function evolution by using an evolutionary search algorithm to automatically find optimal combination of loss functions capturing many (self-supervised) tasks and modalities. Thirdly, we propose an unsupervised representation evaluation metric using distribution matching to a large unlabeled dataset as a prior constraint, based on Zipf's law. This unsupervised constraint, which is not guided by any labeling, produces similar results to weakly-supervised, task-specific ones. The proposed unsupervised representation learning results in a single RGB network and outperforms previous methods. Notably, it is also more effective than several label-based methods (e.g., ImageNet), with the exception of large, fully labeled video datasets. A. J. Piergiovanni, Anelia Angelova, Michael S. Ryoo |
CVPR | 3 |
| 2020 | Password-Conditioned Anonymization and Deanonymization with Face Identity Transformers
Xiuye Gu, Weixin Luo, Michael S. Ryoo, Yong Jae Lee |
ECCV (23) | 3 |
| 2020 | Adversarial Generative Grammars for Human Activity Prediction
A. J. Piergiovanni, Anelia Angelova, Alexander Toshev, Michael S. Ryoo |
ECCV (2) | 4 |
| 2020 | AssembleNet++: Assembling Modality Representations via Attention Connections
Michael S. Ryoo, A. J. Piergiovanni, Juhana Kangaspunta, Anelia Angelova |
ECCV (20) | 1 |
| 2020 | AttentionNAS: Spatiotemporal Attention Cell Search for Video Classification
Xuehan Xiong, Maxim Neumann, A. J. Piergiovanni, Michael S. Ryoo, Anelia Angelova, Kris Makoto Kitani |
ECCV (8) | 5 |
| 2020 | AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures
Michael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, Anelia Angelova |
ICLR | 1 |
| 2020 | AViD Dataset: Anonymized Videos from Diverse CountriesabstractWe introduce a new public video dataset for action recognition: Anonymized Videos from Diverse countries (AViD). Unlike existing public video datasets, AViD is a collection of action videos from many different countries. The motivation is to create a public dataset that would benefit training and pretraining of action recognition models for everybody, rather than making it useful for limited countries. Further, all the face identities in the AViD videos are properly anonymized to protect their privacy. It also is a static dataset where each video is licensed with the creative commons license. We confirm that most of the existing video datasets are statistically biased to only capture action videos from a limited number of countries. We experimentally illustrate that models trained with such biased datasets do not transfer perfectly to action videos from the other countries, and show that AViD addresses such problem. We also confirm that the new AViD dataset could serve as a good dataset for pretraining the models, performing comparably or better than prior datasets. The dataset is available at https://github.com/piergiaj/AViD A. J. Piergiovanni, Michael S. Ryoo |
NeurIPS | 2 |
| 2020 | Learning Multimodal Representations for Unseen ActivitiesabstractWe present a method to learn a joint multimodal representation space that enables recognition of unseen activities in videos. We first compare the effect of placing various constraints on the embedding space using paired text and video data. We also propose a method to improve the joint embedding space using an adversarial formulation, allowing it to benefit from unpaired text and video data. By using unpaired text data, we show the ability to learn a representation that better captures unseen activities. In addition to testing on publicly available datasets, we introduce a new, large-scale text/video dataset. We experimentally confirm that using paired and unpaired data to learn a shared embedding space benefits three difficult tasks (i) zero-shot activity classification, (ii) unsupervised activity discovery, and (iii) unseen activity captioning, outperforming the state-of-the-arts. A. J. Piergiovanni, Michael S. Ryoo |
WACV | 2 |
| 2020 | Model-Based Robot Imitation with Future Image Similarity
A. J. Piergiovanni, Michael S. Ryoo |
Int. J. Comput. Vis. | 3 |
| 2020 | Correction to: Model-Based Robot Imitation with Future Image Similarity
A. J. Piergiovanni, Michael S. Ryoo |
Int. J. Comput. Vis. | 3 |
| 2020 | Toward Collaborative Inferencing of Deep Neural Networks on Internet-of-Things DevicesabstractRecent advancements in deep neural networks (DNNs) have enabled us to solve traditionally challenging problems. To deploy a service based on DNNs, since DNNs are compute intensive, consumers need to rely on compute resources in the cloud. This approach, in addition to creating a dependency on the high-quality network infrastructure and data centers, raises new privacy concerns because of the sharing of private data. These concerns and challenges limit the widespread use of DNN-based applications, so many researchers and companies are trying to optimize DNNs for fast in-the-edge execution. Executing DNNs is further pushed to the edge with the widespread use of embedded processors and ubiquitous wireless networks in Internet-of-Things (IoT) devices. However, inadequate power and computing resources of edge devices, along with the small number of local requests, limit the use of prevalent optimization techniques such as batch processing. In this article, we enable the utilization of the aggregated computing power of several IoT devices by creating a local collaborative network for a subset of DNNs, visual-based applications. In this approach, IoT devices cooperate to conduct single-batch inferencing in real time while exploiting several new model-parallelism methods, which will be introduced in this article. Our approach enhances the collaborative system by creating a balanced and distributed processing pipeline while adjusting the tasks in real time. For experiments, we deploy a system with up to 10 Raspberry Pis and execute state-of-the-art visual models, such as AlexNet, VGG16, Xception, and C3D. Ramyad Hadidi, Jiashen Cao, Michael S. Ryoo, Hyesoon Kim |
IEEE Internet Things J. | 3 |
| 2019 | Representation Flow for Action RecognitionabstractIn this paper, we propose a convolutional layer inspired by optical flow algorithms to learn motion representations. Our representation flow layer is a fully-differentiable layer designed to capture the `flow' of any representation channel within a convolutional neural network for action recognition. Its parameters for iterative flow optimization are learned in an end-to-end fashion together with the other CNN model parameters, maximizing the action recognition performance. Furthermore, we newly introduce the concept of learning `flow of flow' representations by stacking multiple representation flow layers. We conducted extensive experimental evaluations, confirming its advantages over previous recognition models using traditional optical flows in both computational speed and performance. The code is publicly available. A. J. Piergiovanni, Michael S. Ryoo |
CVPR | 2 |
| 2019 | Robustly Executing DNNs in IoT Systems Using Coded Distributed ComputingabstractInternet of Things (IoT) devices have access to an abundance of raw data for processing. With deep neural networks (DNNs), not only the demand for the computing power of IoT devices is increasing, but also privacy concerns are motivating the importance of close-to-edge computation. DNN execution by distributing its computation is common in IoT systems. However, managing unstable latencies in a network and intermittent failures are serious challenges. Our work provides robustness and close-to-zero recovery latency by adapting coded distributed computing (CDC). We analyze robust execution on a mesh of Raspberry Pis by studying four DNNs. Ramyad Hadidi, Jiashen Cao, Michael S. Ryoo, Hyesoon Kim |
DAC | 3 |
| 2019 | Evolving Space-Time Neural Architectures for VideosabstractWe present a new method for finding video CNN architectures that more optimally capture rich spatio-temporal information in videos. Previous work, taking advantage of 3D convolutions, obtained promising results by manually designing CNN video architectures. We here develop a novel evolutionary algorithm that automatically explores models with different types and combinations of layers to jointly learn interactions between spatial and temporal aspects of video representations. We demonstrate the generality of this algorithm by applying it to two meta-architectures. Further, we propose a new component, the iTGM layer, which more efficiently utilizes its parameters to allow learning of space-time interactions over longer time horizons. The iTGM layer is often preferred by the evolutionary algorithm and allows building cost-efficient networks. The proposed approach discovers new diverse and interesting video architectures that were unknown previously. More importantly they are both more accurate and faster than prior models, and outperform the state-of-the-art results on four datasets: Kinetics, Charades, Moments in Time and HMDB. We will open source the code and models, to encourage future model development. A. J. Piergiovanni, Anelia Angelova, Alexander Toshev, Michael S. Ryoo |
ICCV | 4 |
| 2019 | Temporal Gaussian Mixture Layer for VideosabstractWe introduce a new convolutional layer named the Temporal Gaussian Mixture (TGM) layer and present how it can be used to efficiently capture longer-term temporal information in continuous activity videos. The TGM layer is a temporal convolutional layer governed by a much smaller set of parameters (e.g., location/variance of Gaussians) that are fully differentiable. We present our fully convolutional video models with multiple TGM layers for activity detection. The extensive experiments on multiple datasets, including Charades and MultiTHUMOS, confirm the effectiveness of TGM layers, significantly outperforming the state-of-the-arts. A. J. Piergiovanni, Michael S. Ryoo |
ICML | 2 |
| 2019 | Privacy-Preserving Robot Vision with Anonymized Faces by Extreme Low ResolutionabstractAs smart cameras are becoming ubiquitous in mobile robot systems, there is an increasing concern in camera devices invading people's privacy by recording unwanted images. We want to fundamentally protect privacy by blurring unwanted blocks in images, such as faces, yet ensure that the robots can understand the video for their perception. In this paper, we propose a novel mobile robot framework with a deep learning-based privacy-preserving camera system. The proposed camera system detects privacy-sensitive blocks, i.e., human face, from extreme low resolution (LR) images, and then dynamically enhances the resolution of only privacy-insensitive blocks, e.g., backgrounds. Keeping all the face blocks to be extreme LR of 15x15 pixels, we can guarantee that human faces are never at high resolution (HR) in any of processing or memory, thus yielding strong privacy protection even from cracking or backdoors. Our camera system produces an image on a real-time basis, the human faces of which are in extreme LR while the backgrounds are in HR. We experimentally confirm that our proposed face detection camera system outperforms the state-of-the-art small face detection algorithm, while the robot performs ORB-SLAM2 well even with videos of extreme LR faces. Therefore, with the proposed system, we do not too much sacrifice robot perception performance to protect privacy. Myeung Un Kim, Harim Lee, Hyun Jong Yang, Michael S. Ryoo |
IROS | 4 |
| 2019 | Learning Real-World Robot Policies by DreamingabstractLearning to control robots directly based on images is a primary challenge in robotics. However, many existing reinforcement learning approaches require iteratively obtaining millions of robot samples to learn a policy, which can take significant time. In this paper, we focus on learning a realistic world model capturing the dynamics of scene changes conditioned on robot actions. Our dreaming model can emulate samples equivalent to a sequence of images from the actual environment, technically by learning an action-conditioned future representation/scene regressor. This allows the agent to learn action policies (i.e., visuomotor policies) by interacting with the dreaming model rather than the real-world. We experimentally confirm that our dreaming model enables robot learning of policies that transfer to the real-world. A. J. Piergiovanni, Michael S. Ryoo |
IROS | 3 |
| 2018 | Extreme Low Resolution Activity Recognition With Multi-Siamese Embedding LearningabstractThis paper presents an approach for recognizing human activities from extreme low resolution (e.g., 16x12) videos. Extreme low resolution recognition is not only necessary for analyzing actions at a distance but also is crucial for enabling privacy-preserving recognition of human activities. We design a new two-stream multi-Siamese convolutional neural network. The idea is to explicitly capture the inherent property of low resolution (LR) videos that two images originated from the exact same scene often have totally different pixel values depending on their LR transformations. Our approach learns the shared embedding space that maps LR videos with the same content to the same location regardless of their transformations. We experimentally confirm that our approach of jointly learning such transform robust LR video representation and the classifier outperforms the previous state-of-the-art low resolution recognition approaches on two public standard datasets by a meaningful margin. Michael S. Ryoo, Kiyoon Kim, Hyun Jong Yang |
AAAI | 1 |
| 2018 | Learning Latent Super-Events to Detect Multiple Activities in VideosabstractIn this paper, we introduce the concept of learning latent super-events from activity videos, and present how it benefits activity detection in continuous videos. We define a super-event as a set of multiple events occurring together in videos with a particular temporal organization; it is the opposite concept of sub-events. Real-world videos contain multiple activities and are rarely segmented (e.g., surveillance videos), and learning latent super-events allows the model to capture how the events are temporally related in videos. We design temporal structure filters that enable the model to focus on particular sub-intervals of the videos, and use them together with a soft attention mechanism to learn representations of latent super-events. Super-event representations are combined with per-frame or per-segment CNNs to provide frame-level annotations. Our approach is designed to be fully differentiable, enabling end-to-end learning of latent super-event representations jointly with the activity detector using them. Our experiments with multiple public video datasets confirm that the proposed concept of latent super-event learning significantly benefits activity detection, advancing the state-of-the-arts. A. J. Piergiovanni, Michael S. Ryoo |
CVPR | 2 |
| 2018 | Learning to Anonymize Faces for Privacy Preserving Action Detection
Zhongzheng Ren, Yong Jae Lee, Michael S. Ryoo |
ECCV (1) | 3 |
| 2018 | Joint Person Segmentation and Identification in Synchronized First- and Third-Person Videos
Chenyou Fan, Michael S. Ryoo, David Crandall |
ECCV (1) | 4 |
| 2017 | Title Learning Latent Subevents in Activity Videos Using Temporal Attention FiltersabstractIn this paper, we newly introduce the concept of temporal attention filters, and describe how they can be used for human activity recognition from videos. Many high-level activities are often composed of multiple temporal parts (e.g., sub-events) with different duration/speed, and our objective is to make the model explicitly learn such temporal structure using multiple attention filters and benefit from them. Our temporal filters are designed to be fully differentiable, allowing end-of-end training of the temporal filters together with the underlying frame-based or segment-based convolutional neural network architectures. This paper presents an approach of learning a set of optimal static temporal attention filters to be shared across different videos, and extends this approach to dynamically adjust attention filters per testing video using recurrent long short-term memory networks (LSTMs). This allows our temporal attention filters to learn latent sub-events specific to each activity. We experimentally confirm that the proposed concept of temporal attention filters benefits the activity recognition, and we visualize the learned latent sub-events. A. J. Piergiovanni, Chenyou Fan, Michael S. Ryoo |
AAAI | 3 |
| 2017 | Privacy-Preserving Human Activity Recognition from Extreme Low ResolutionabstractPrivacy protection from surreptitious video recordings is an important societal challenge. We desire a computer vision system (e.g., a robot) that can recognize human activities and assist our daily life, yet ensure that it is not recording video that may invade our privacy. This paper presents a fundamental approach to address such contradicting objectives: human activity recognition while only using extreme low-resolution (e.g., 16x12) anonymized videos. We introduce the paradigm of inverse super resolution (ISR), the concept of learning the optimal set of image transformations to generate multiple low-resolution (LR) training videos from a single video. Our ISR learns different types of sub-pixel transformations optimized for the activity classification, allowing the classifier to best take advantage of existing high-resolution videos (e.g., YouTube videos) by creating multiple LR training videos tailored for the problem. We experimentally confirm that the paradigm of inverse super resolution is able to benefit activity recognition from extreme low-resolution videos. Michael S. Ryoo, Brandon Rothrock, Charles Fleming, Hyun Jong Yang |
AAAI | 1 |
| 2017 | Identifying First-Person Camera Wearers in Third-Person VideosabstractWe consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establish person-level correspondences across first-and third-person videos, which is challenging because the camera wearer is not visible from his/her own egocentric video, preventing the use of direct feature matching. In this paper, we propose a new semi-Siamese Convolutional Neural Network architecture to address this novel challenge. We formulate the problem as learning a joint embedding space for first-and third-person videos that considers both spatial-and motion-domain cues. A new triplet loss function is designed to minimize the distance between correct first-and third-person matches while maximizing the distance between incorrect ones. This end-to-end approach performs significantly better than several baselines, in part by learning the first-and third-person features optimized for matching jointly with the distance measure itself. Chenyou Fan, Jangwon Lee 0002, Krishna Kumar Singh, Yong Jae Lee, David Crandall, Michael S. Ryoo |
CVPR | 7 |
| 2017 | Learning social affordance grammar from videos: Transferring human interactions to human-robot interactionsabstractIn this paper, we present a general framework for learning social affordance grammar as a spatiotemporal AND-OR graph (ST-AOG) from RGB-D videos of human interactions, and transfer the grammar to humanoids to enable a real-time motion inference for human-robot interaction (HRI). Based on Gibbs sampling, our weakly supervised grammar learning can automatically construct a hierarchical representation of an interaction with long-term joint sub-tasks of both agents and short term atomic actions of individual agents. Based on a new RGB-D video dataset with rich instances of human interactions, our experiments of Baxter simulation, human evaluation, and real Baxter test demonstrate that the model learned from limited training data successfully generates human-like behaviors in unseen scenarios and outperforms both baselines. Tianmin Shu, Xiaofeng Gao 0002, Michael S. Ryoo, Song-Chun Zhu |
ICRA | 3 |
| 2017 | Multi-Type Activity Recognition from a Robot's ViewpointabstractThe literature in computer vision is rich of works where different types of activities -- single actions, two persons interactions or ego-centric activities, to name a few -- have been analyzed. However, traditional methods treat such types of activities separately, while in real settings detecting and recognizing different types of activities simultaneously is necessary. We first design a new unified descriptor, called Relation History Image (RHI), which can be extracted from all the activity types we are interested in. We then formulate an optimization procedure to detect and recognize activities of different types. We assess our approach on a new dataset recorded from a robot-centric perspective as well as on publicly available datasets, and evaluate its quality compared to multiple baselines. Ilaria Gori, Jake K. Aggarwal, Larry H. Matthies, Michael S. Ryoo |
IJCAI | 4 |
| 2017 | Learning robot activities from first-person human videos using convolutional future regressionabstractWe design a new approach that allows robot learning of new activities from unlabeled human example videos. Given videos of humans executing the same activity from a human's viewpoint (i.e., first-person videos), our objective is to make the robot learn the temporal structure of the activity as its future regression network, and learn to transfer such model for its own motor execution. We present a new deep learning model: We extend the state-of-the-art convolutional object detection network for the representation/estimation of human hands in training videos, and newly introduce the concept of using a fully convolutional network to regress (i.e., predict) the intermediate scene representation corresponding to the future frame (e.g., 1-2 seconds later). Combining these allows direct prediction of future locations of human hands and objects, which enables the robot to infer the motor control plan using our manipulation network. We experimentally confirm that our approach makes learning of robot activities from unlabeled human interaction videos possible, and demonstrate that our robot is able to execute the learned collaborative activities in real-time directly based on its camera input. Jangwon Lee 0002, Michael S. Ryoo |
IROS | 2 |
| 2016 | Learning Social Affordance for Human-Robot Interaction
Tianmin Shu, Michael S. Ryoo, Song-Chun Zhu |
IJCAI | 2 |
| 2016 | First-Person Activity Recognition: Feature, Temporal Structure, and Prediction
Michael S. Ryoo, Larry H. Matthies |
Int. J. Comput. Vis. | 1 |
| 2015 | Pooled motion features for first-person videosabstractIn this paper, we present a new feature representation for first-person videos. In first-person video understanding (e.g., activity recognition), it is very important to capture both entire scene dynamics (i.e., egomotion) and salient local motion observed in videos. We describe a representation framework based on time series pooling, which is designed to ab] short-term/long-term changes in feature descriptor elements. The idea is to keep track of how descriptor values are changing over time and summarize them to represent motion in the activity video. The framework is general, handling any types of per-frame feature descriptors including conventional motion descriptors like histogram of optical flows (HOF) as well as appearance descriptors from more recent convolutional neural networks (CNN). We experimentally confirm that our approach clearly outperforms previous feature representations including bag-of-visual-words and improved Fisher vector (IFV) when using identical underlying feature descriptors. We also confirm that our feature representation has superior performance to existing state-of-the-art features like local spatio-temporal features and Improved Trajectory Features (originally developed for 3rd-person videos) when handling first-person videos. Multiple first-person activity datasets were tested under various settings to confirm these findings. Michael S. Ryoo, Brandon Rothrock, Larry H. Matthies |
CVPR | 1 |
| 2015 | Robot-Centric Activity Prediction from First-Person Videos: What Will They Do to Me'abstractIn this paper, we present a core technology to enable robot recognition of human activities during human-robot interactions. In particular, we propose a methodology for early recognition of activities from robot-centric videos (i.e., first-person videos) obtained from a robot's viewpoint during its interaction with humans. Early recognition, which is also known as activity prediction, is an ability to infer an ongoing activity at its early stage. We present an algorithm to recognize human activities targeting the camera from streaming videos, enabling the robot to predict intended activities of the interacting person as early as possible and take fast reactions to such activities (e.g., avoiding harmful events targeting itself before they actually occur). We introduce the novel concept of 'onset' that efficiently summarizes pre-activity observations, and design a recognition approach to consider event history in addition to visual features from first-person videos. We propose to represent an onset using a cascade histogram of time series gradients, and we describe a novel algorithmic setup to take advantage of such onset for early recognition of activities. The experimental results clearly illustrate that the proposed concept of onset enables better/earlier recognition of human activities from first-person videos collected with a robot. Michael S. Ryoo, Thomas J. Fuchs, Jake K. Aggarwal, Larry H. Matthies |
HRI | 1 |
| 2015 | Robot-centric Activity Recognition from First-Person RGB-D VideosabstractWe present a framework and algorithm to analyze first person RGBD videos captured from the robot while physically interacting with humans. Specifically, we explore reactions and interactions of persons facing a mobile robot from a robot centric view. This new perspective offers social awareness to the robots, enabling interesting applications. As far as we know, there is no public 3D dataset for this problem. Therefore, we record two multi-modal first-person RGBD datasets that reflect the setting we are analyzing. We use a humanoid and a non-humanoid robot equipped with a Kinect. Notably, the videos contain a high percentage of ego-motion due to the robot self-exploration as well as its reactions to the persons' interactions. We show that separating the descriptors extracted from ego-motion and independent motion areas, and using them both, allows us to achieve superior recognition results. Experiments show that our algorithm recognizes the activities effectively and outperforms other state-of-the-art methods on related tasks. Ilaria Gori, Jake K. Aggarwal, Michael S. Ryoo |
WACV | 4 |
| 2014 | First-Person Animal Activity Recognition from Egocentric VideosabstractThis paper introduces the concept of first-person animal activity recognition, the problem of recognizing activities from a view-point of an animal (e.g., a dog). Similar to first-person activity recognition scenarios where humans wear cameras, our approach estimates activities performed by an animal wearing a camera. This enables monitoring and understanding of natural animal behaviors even when there are no people around them. Its applications include automated logging of animal behaviors for medical/biology experiments, monitoring of pets, and investigation of wildlife patterns. In this paper, we construct a new dataset composed of first-person animal videos obtained by mounting a camera on each of the four pet dogs. Our new dataset consists of 10 activities containing a heavy/fair amount of ego-motion. We implemented multiple baseline approaches to recognize activities from such videos while utilizing multiple types of global/local motion features. Animal ego-actions as well as human-animal interactions are recognized with the baseline approaches, and we discuss experimental results. Yumi Iwashita, Asamichi Takamine, Ryo Kurazume, Michael S. Ryoo |
ICPR | 4 |
| 2013 | Recognizing Humans in Motion: Trajectory-based Aerial Video AnalysisabstractWe propose a novel method for recognizing people in aerial surveillance videos. Aerial surveillance images cover a wide area at low resolution. In order to detect objects (e.g., pedestrians) from such videos, conventional methods either utilize appearance information from raw videos or extract blob information from background subtraction results. However, people seen in low resolution images have less appearance information, and hence are very difficulty to classify based on their appearance or blob size. In addition, due to heavy camera movements caused by aerial vehicle ego-motion and wind, the system is expected to generate many noisy false detections including parallax. The idea presented in this paper is to detect and classify objects from aerial videos based on their motion: we analyze a trajectory of each object candidate, deciding whether it is a person-of-interest or simple noise based on how it moved. After objects are tracked by a Kalman filter-based tracking, we represent their motion as multi-scale histograms of ‘orientation changes’, which efficiently captures movements displayed by objects. Random forest classifiers are applied to our new representation to make the decision. The experimental results illustrate that our approach recognizes objects-of-interest (i.e., humans) even when there exist a large number of false detection/tracking, and it does it more reliably compared to the approaches with previous paradigm. Yumi Iwashita, Michael S. Ryoo, Thomas J. Fuchs, Curtis Padgett |
BMVC | 2 |
| 2013 | First-Person Activity Recognition: What Are They Doing to Me?abstractThis paper discusses the problem of recognizing interaction-level human activities from a first-person viewpoint. The goal is to enable an observer (e.g., a robot or a wearable camera) to understand 'what activity others are performing to it' from continuous video inputs. These include friendly interactions such as 'a person hugging the observer' as well as hostile interactions like 'punching the observer' or 'throwing objects to the observer', whose videos involve a large amount of camera ego-motion caused by physical interactions. The paper investigates multi-channel kernels to integrate global and local motion information, and presents a new activity learning/recognition methodology that explicitly considers temporal structures displayed in first-person activity videos. In our experiments, we not only show classification results with segmented videos, but also confirm that our new approach is able to detect activities from continuous videos reliably. Michael S. Ryoo, Larry H. Matthies |
CVPR | 1 |
| 2013 | Personal driving diary: Automated recognition of driving events from first-person videos
Michael S. Ryoo, Sunglok Choi, Ji Hoon Joung, Jae-Yeong Lee, Wonpil Yu |
Comput. Vis. Image Underst. | 1 |
| 2012 | Reliable object detection and segmentation using inpaintingabstractThis paper presents a novel object detection and segmentation method utilizing an inpainting algorithm. Inpainting is a concept of recovering missing image regions based on their surroundings, which were originally used for restoration of damaged paintings. In this paper, we newly utilize inpainting to judge whether an object candidate region includes the foreground object or not. The key idea is that if we erase a certain region from an image, the inpainting algorithm is expected to recover the erased image only when it belongs a background area (i.e. only when there is no object in it). By measuring the similarity between the inpainted region and the original image region, our approach filters out false detections while maintaining true object detections. Furthermore, we take advantage of the inpainting for object segmentation, since our approach is designed to explicitly distinguish foreground areas from its background. Experimental results confirm that our approach applied to baseline detectors enables better recognition of objects, obtaining higher accuracies. We illustrate how our inpainting-based detection/segmentation approach benefits the object detection using two different pedestrian datasets. Ji Hoon Joung, Michael S. Ryoo, Sunglok Choi, Sung-Rak Kim |
IROS | 2 |
| 2012 | Toward a unified framework of motion understanding
Jake K. Aggarwal, Michael S. Ryoo |
Image Vis. Comput. | 2 |
| 2011 | Human activity prediction: Early recognition of ongoing activities from streaming videosabstractIn this paper, we present a novel approach of human activity prediction. Human activity prediction is a probabilistic process of inferring ongoing activities from videos only containing onsets (i.e. the beginning part) of the activities. The goal is to enable early recognition of unfinished activities as opposed to the after-the-fact classification of completed activities. Activity prediction methodologies are particularly necessary for surveillance systems which are required to prevent crimes and dangerous activities from occurring. We probabilistically formulate the activity prediction problem, and introduce new methodologies designed for the prediction. We represent an activity as an integral histogram of spatio-temporal features, efficiently modeling how feature distributions change over time. The new recognition methodology named dynamic bag-of-words is developed, which considers sequential nature of human activities while maintaining advantages of the bag-of-words to handle noisy observations. Our experiments confirm that our approach reliably recognizes ongoing activities from streaming videos with a high accuracy. Michael S. Ryoo |
ICCV | 1 |
| 2011 | Personal driving diary: Constructing a video archive of everyday driving eventsabstractIn this paper, we introduce the concept of personal driving diary. A personal driving diary is a multimedia archive of a person's daily driving experience, describing important driving events of the user with annotated videos. This paper presents an automated system that constructs such multimedia diary by analyzing videos obtained from a vehicle-mounted camera. The proposed system recognizes important interactions between the driving vehicle and the others from videos (e.g. accident, overtaking, ...), and labels them together with its contextual knowledge on the vehicle (e.g. its physical location on the map) to construct an event log. A novel decision tree based activity recognizer that incrementally learns driving events from first-person view videos is designed. The constructed diary enables efficient searching and event-based browsing of video clips, which helps the user to retrieve videos of dangerous situations and analyze his/her driving habits statistically. Our experiment confirms that the proposed system reliably generates driving diaries by annotating learned vehicle events. Michael S. Ryoo, Jae-Yeong Lee, Ji Hoon Joung, Sunglok Choi, Wonpil Yu |
WACV | 1 |
| 2011 | One video is sufficient? Human activity recognition using active video compositionabstractIn this paper, we present a novel human activity recognition approach that only requires a single video example per activity. We introduce the paradigm of active video composition, which enables one-example recognition of complex activities. The idea is to automatically create a large number of semi-artificial training videos called composed videos by manipulating an original human activity video. A methodology to automatically compose activity videos having different backgrounds, translations, scales, actors, and movement structures is described in this paper. Furthermore, an active learning algorithm to model the temporal structure of the human activity has been designed, preventing the generation of composed training videos violating the structural constraints of the activity. The intention is to generate composed videos having correct organizations, and take advantage of them for the training of the recognition system. In contrast to previous passive recognition systems relying only on given training videos, our methodology actively composes necessary training videos that the system is expected to observe in its environment. Experimental results illustrate that a single fully labeled video per activity is sufficient for our methodology to reliably recognize human activities by utilizing composed training videos. Michael S. Ryoo, Wonpil Yu |
WACV | 1 |
| 2011 | Stochastic Representation and Recognition of High-Level Group Activities
Michael S. Ryoo, Jake K. Aggarwal |
Int. J. Comput. Vis. | 1 |
| 2010 | A task-driven intelligent workspace system to provide guidance feedback
Michael S. Ryoo, Kristen Grauman, Jake K. Aggarwal |
Comput. Vis. Image Underst. | 1 |
| 2009 | Spatio-temporal relationship match: Video structure comparison for recognition of complex human activitiesabstractHuman activity recognition is a challenging task, especially when its background is unknown or changing, and when scale or illumination differs in each video. Approaches utilizing spatio-temporal local features have proved that they are able to cope with such difficulties, but they mainly focused on classifying short videos of simple periodic actions. In this paper, we present a new activity recognition methodology that overcomes the limitations of the previous approaches using local features. We introduce a novel matching, spatio-temporal relationship match, which is designed to measure structural similarity between sets of features extracted from two videos. Our match hierarchically considers spatio-temporal relationships among feature points, thereby enabling detection and localization of complex non-periodic activities. In contrast to previous approaches to `classify' videos, our approach is designed to `detect and localize' all occurring activities from continuous videos where multiple actors and pedestrians are present. We implement and test our methodology on a newly-introduced dataset containing videos of multiple interacting persons and individual pedestrians. The results confirm that our system is able to recognize complex non-periodic activities (e.g. `push' and `hug') from sets of spatio-temporal features even when multiple activities are present in the scene. Michael S. Ryoo, Jake K. Aggarwal |
ICCV | 1 |
| 2009 | Semantic Representation and Recognition of Continued and Recursive Human Activities
Michael S. Ryoo, Jake K. Aggarwal |
Int. J. Comput. Vis. | 1 |
| 2009 | Detection of object abandonment using temporal logic
Medha Bhargava, Chia-Chih Chen, Michael S. Ryoo, Jake K. Aggarwal |
Mach. Vis. Appl. | 3 |
| 2009 | Real-Time Illegal Parking Detection in Outdoor Environments Using 1-D TransformationabstractWith decreasing costs of high-quality surveillance systems, human activity detection and tracking has become increasingly practical. Accordingly, automated systems have been designed for numerous detection tasks, but the task of detecting illegally parked vehicles has been left largely to the human operators of surveillance systems. We propose a methodology for detecting this event in real time by applying a novel image projection that reduces the dimensionality of the data and, thus, reduces the computational complexity of the segmentation and tracking processes. After event detection, we invert the transformation to recover the original appearance of the vehicle and to allow for further processing that may require 2-D data. We evaluate the performance of our algorithm using the i-LIDS vehicle detection challenge datasets as well as videos we have taken ourselves. These videos test the algorithm in a variety of outdoor conditions, including nighttime video and instances of sudden changes in weather. Jong Taek Lee, Michael S. Ryoo, Matthew Riley, Jake K. Aggarwal |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Observe-and-explain: A new approach for multiple hypotheses tracking of humans and objectsabstractThis paper presents a novel approach for tracking humans and objects under severe occlusion. We introduce a new paradigm for multiple hypotheses tracking, observe-and-explain, as opposed to the previous paradigm of hypothesize-and-test. Our approach efficiently enumerates multiple possibilities of tracking by generating several likely dasiaexplanationspsila after concatenating a sufficient amount of observations. The computational advantages of our approach over the previous paradigm under severe occlusions are presented. The tracking system is implemented and tested using the i-Lids dataset, which consists of videos of humans and objects moving in a London subway station. The experimental results show that our new approach is able to track humans and objects accurately and reliably even when they are completely occluded, illustrating its advantage over previous approaches. Michael S. Ryoo, Jake K. Aggarwal |
CVPR | 1 |
| 2008 | Human activities: Handling uncertainties using fuzzy time intervalsabstractPersons may perform an activity in many different styles, or noise may cause an identical activity to have different temporal structures. We present a robust methodology for recognition of such human activities. The recognition approach presented in this paper is able to handle person-dependent and situation-dependent uncertainties and variations of human activity executions. Our system reliably recognizes human activities with such execution variations, by semantically measuring the similarity between the observations generated by an activity execution and its optimal structure. The system detects fuzzy time intervals associated with low-level gestures of a person, and matches them hierarchically with the representation of the activity that the system is maintaining. Our system is tested for eight types of simple human interactions such as `pushing¿ and `shaking hands¿, as well as complex recursive interactions like `fighting¿ and `greeting¿. The results show that the performance of our system is superior to that of the previous systems using deterministic time intervals. Michael S. Ryoo, Jake K. Aggarwal |
ICPR | 1 |
| 2007 | Detection of abandoned objects in crowded environmentsabstractWith concerns about terrorism and global security on the rise, it has become vital to have in place efficient threat detection systems that can detect and recognize potentially dangerous situations, and alert the authorities to take appropriate action. Of particular significance is the case of unattended objects in mass transit areas. This paper describes a general framework that recognizes the event of someone leaving a piece of baggage unattended in forbidden areas. Our approach involves the recognition of four sub-events that characterize the activity of interest. When an unaccompanied bag is detected, the system analyzes its history to determine its most likely owner(s), where the owner is defined as the person who brought the bag into the scene before leaving it unattended. Through subsequent frames, the system keeps a lookout for the owner, whose presence in or disappearance from the scene defines the status of the bag, and decides the appropriate course of action. The system was successfully tested on the i-LIDS dataset. Medha Bhargava, Chia-Chih Chen, Michael S. Ryoo, Jake K. Aggarwal |
AVSS | 3 |
| 2007 | Real-time detection of illegally parked vehicles using 1-D transformationabstractWith decreasing costs of high quality surveillance systems, human activity detection and tracking has become increasingly practical. Accordingly, automated systems have been designed for numerous detection tasks, but the task of detecting illegally parked vehicles has been left largely to the human operators of surveillance systems. We propose a methodology for detecting this event in realtime by applying a novel image projection that reduces the dimensionality of the image data and thus reduces the computational complexity of the segmentation and tracking processes. After event detection, we invert the transformation to recover the original appearance of the vehicle and to allow for further processing that may require the two dimensional data. The proposed algorithm is able to successfully recognize illegally parked vehicles in real-time in the i-LIDS bag and vehicle detection challenge datasets. Jong Taek Lee, Michael S. Ryoo, Matthew Riley, Jake K. Aggarwal |
AVSS | 2 |
| 2007 | Hierarchical Recognition of Human Activities Interacting with ObjectsabstractThe paper presents a system that recognizes humans interacting with objects. We delineate a new framework that integrates object recognition, motion estimation, and semantic-level recognition for the reliable recognition of hierarchical human-object interactions. The framework is designed to integrate recognition decisions made by each component, and to probabilistically compensate for the failure of the components with the use of the decisions made by the other components. As a result, human-object interactions in an airport-like environment, such as 'a person carrying a baggage', 'a person leaving his/her baggage', or 'a person snatching another's baggage', are recognized. The experimental results show that not only the performance of the final activity recognition is superior to that of previous approaches, but also the accuracy of the object recognition and the motion estimation increases using feedback from the semantic layer. Several real examples illustrate the superior performance in recognition and semantic description of occurring events. Michael S. Ryoo, Jake K. Aggarwal |
CVPR | 1 |
| 2007 | Robust Human-Computer Interaction System Guiding a User by Providing Feedback
Michael S. Ryoo, Jake K. Aggarwal |
IJCAI | 1 |
| 2006 | Recognition of Composite Human Activities through Context-Free Grammar Based RepresentationabstractThis paper describes a general methodology for automated recognition of complex human activities. The methodology uses a context-free grammar (CFG) based representation scheme to represent composite actions and interactions. The CFG-based representation enables us to formally define complex human activities based on simple actions or movements. Human activities are classified into three categories: atomic action, composite action, and interaction. Our system is not only able to represent complex human activities formally, but also able to recognize represented actions and interactions with high accuracy. Image sequences are processed to extract poses and gestures. Based on gestures, the system detects actions and interactions occurring in a sequence of image frames. Our results show that the system is able to represent composite actions and interactions naturally. The system was tested to represent and recognize eight types of interactions: approach, depart, point, shake-hands, hug, punch, kick, and push. The experiments show that the system can recognize sequences of represented composite actions and interactions with a high recognition rate. Michael S. Ryoo, Jake K. Aggarwal |
CVPR (2) | 1 |
| 2005 | Affective Dialogue Communication System with Emotional Memories for Humanoid Robots
Michael S. Ryoo, Yongho Seo, Hye-Won Jung, Hyun Seung Yang |
ACII | 1 |
| 2005 | Evolving neural network ensembles for control problemsabstractIn neuroevolution, a genetic algorithm is used to evolve a neural network to perform a particular task. The standard approach is to evolve a population over a number of generations, and then select the final generation's champion as the end result. However, it is possible that there is valuable information present in the population that is not captured by the champion. The standard approach ignores all such information. One possible solution to this problem is to combine multiple individuals from the final population into an ensemble. This approach has been successful in supervised classification tasks, and in this paper, it is extended to evolutionary reinforcement learning in control problems. The method is evaluated on a challenging extension of the classic pole balancing task, demonstrating that an ensemble can achieve significantly better performance than the champion alone. David Pardoe, Michael S. Ryoo, Risto Miikkulainen |
GECCO | 2 |