Yoichi Sato 0001

dblp:64/5287-1 · DBLP profile ↗
← Back
189ranked-venue papers
7as first author
46since 2021 · last 2026
0000-0003-0097-4537ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 134 · 3 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 122 · 7 first-author · 34 since 2021Human-computer interaction and ubiquitous computing · 26 · 2 first-author · 2 since 2021Systems, architecture and hardware · 6Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Leveraging RGB Images for Pre-training of Event-Based Hand Pose Estimation
Ruicong Liu, Takehiko Ohkawa, Tze Ho Elden Tse, Mingfang Zhang 0002, Angela Yao, Yoichi Sato 0001
ICPR (8)6
2025 Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities
abstract
We address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze signals. While traditional methods rely on manually labeled mistakes, our approach does not require mistake annotations, hence overcoming the need of domain-specific labeled data. Based on the observation that eye movements closely follow object manipulation activities, we assess to what extent eye-gaze signals can support mistake detection, proposing to identify deviations in attention patterns measured through a gaze tracker with respect to those estimated by a gaze prediction model. Since predicting gaze in video is characterized by high uncertainty, we propose a novel gaze completion task, where eye fixations are predicted from visual observations and partial gaze trajectories, and contribute a novel gaze completion approach which explicitly models correlations between gaze information and local visual tokens. Inconsistencies between predicted and observed gaze trajectories act as an indicator to identify mistakes. Experiments highlight the effectiveness of the proposed approach in different settings, with relative gains up to +14%, +11%, and +5% in EPIC-Tent, HoloAssist and IndustReal respectively, remarkably matching results of supervised approaches without seeing any labels. We further show that gaze-based analysis is particularly useful in the presence of skilled actions, low action execution confidence, and actions requiring hand-eye coordination and object manipulation skills. Our method is ranked first on the HoloAssist Mistake Detection challenge.
Michele Mazzamuto, Antonino Furnari, Yoichi Sato 0001, Giovanni Maria Farinella
CVPR3
2025 ChartQC: Question Classification from Human Attention Data on Charts
Takumi Nishiyasu, Tobias Kostorz, Yao Wang 0018, Yoichi Sato 0001, Andreas Bulling
ETRA4
2025 Generative Modeling of Shape-Dependent Self-Contact Human Poses
Takehiko Ohkawa, Shunsuke Saito, Jason M. Saragih, Fabian Prada, Shoou-I Yu, Ryosuke Furuta, Yoichi Sato 0001, Takaaki Shiratori
ICCV9
2025 Egocentric Action-Aware Inertial Localization in Point Clouds with Vision-Language Guidance
abstract
This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individual within a 3D point cloud. Human inertial localization is challenging due to IMU sensor noise that causes trajectory drift over time. The diversity of human actions further complicates IMU signal processing by introducing various motion patterns. Nevertheless, we observe that some actions captured by the head-mounted IMU correlate with spatial environmental structures (e.g., bending down to look inside an oven, washing dishes next to a sink), thereby serving as spatial anchors to compensate for the localization drift. The proposed EAIL framework learns such correlations via hierarchical multi-modal alignment with vision-language guidance. By assuming that the 3D point cloud of the environment is available, it contrastively learns modality encoders that align short-term egocentric action cues in IMU signals with local environmental features in the point cloud. The learning process is enhanced using concurrently collected vision and language signals to improve multimodal alignment. The learned encoders are then used in reasoning the IMU data and the point cloud over time and space to perform inertial localization. Interestingly, these encoders can further be utilized to recognize the corresponding sequence of actions as a by-product. Extensive experiments demonstrate the effectiveness of the proposed framework over state-of-the-art inertial localization and inertial action recognition baselines.
Mingfang Zhang 0002, Ryo Yonetani, Yifei Huang 0002, Liangyang Ouyang, Ruicong Liu, Yoichi Sato 0001
ICCV6
2025 SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training
abstract
We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SiMHand. Pre-training with large-scale images achieves promising results in various tasks, but prior methods for 3D hand pose pre-training have not fully utilized the potential of diverse hand images accessible from in-the-wild videos. To facilitate scalable pre-training, we first prepare an extensive pool of hand images from in-the-wild videos and design our pre-training method with contrastive learning. Specifically, we collect over 2.0M hand images from recent human-centric videos, such as 100DOH and Ego4D. To extract discriminative information from these images, we focus on the similarity of hands: pairs of non-identical samples with similar hand poses. We then propose a novel contrastive learning method that embeds similar hand pairs closer in the feature space. Our method not only learns from similar samples but also adaptively weights the contrastive learning loss based on inter-sample distance, leading to additional performance gains. Our experiments demonstrate that our method outperforms conventional contrastive learning approaches that produce positive pairs solely from a single image with data augmentation. We achieve significant improvements over the state-of-the-art method (PeCLR) in various datasets, with gains of 15% on FreiHand, 10% on DexYCB, and 4% on AssemblyHands. Our code is available at https://github.com/ut-vision/SiMHand.
Nie Lin, Takehiko Ohkawa, Yifei Huang 0002, Mingfang Zhang 0002, Minjie Cai, Ryosuke Furuta, Yoichi Sato 0001
ICLR8
2025 Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura, Ryosuke Furuta, Atsushi Hashimoto 0001, Yoshitaka Ushiku, Yoichi Sato 0001
WACV7
2025 Learning Multiple Object States from Actions via Large Language Models
abstract
Recognizing the states of objects in a video is crucial in understanding the scene beyond actions and objects. For instance, an egg can be raw, cracked, and whisked while cooking an omelet, and these states can coexist simulta-neously (an egg can be both raw and whisked). However, most existing research assumes single object state change (e.g. uncracked → cracked), overlooking the coexisting nature of multiple object states and the influence of past states on the current state. We formulate object state recognition as a multi-label classification task that explicitly handles multiple states. We then propose to learn multiple object states from narrated videos by leveraging LLMs to generate pseudo-labels from the transcribed narrations, capturing the influence of past states. The challenge is that narrations mostly describe human actions in the video but rarely explain object states. Therefore, we use LLM's knowledge of the relationship between actions and states to derive the missing object states. We further accumulate the derived object states to consider the past state contexts to infer current object state pseudo-labels. We newly collect Multiple Object States Transition (MOST) dataset, which includes manual multi-label annotation for evaluation purposes, covering 60 object states across six object categories. Experimental results show that our model trained on LLM-generated pseudo-labels significantly outperforms strong vision-language models, demonstrating the effectiveness of our pseudo-labeling framework that considers past context via LLMs.
Masatoshi Tateno, Takuma Yagi, Ryosuke Furuta, Yoichi Sato 0001
WACV4
2025 Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
abstract
We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions—including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. https://ego-exo4d-data.org/
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, David Crandall, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray
Int. J. Comput. Vis.97
2025 FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
abstract
Abstract In the development of science, accurate and reproducible documentation of the experimental process is crucial. Automatic recognition of the actions in experiments from videos would help experimenters by complementing the recording of experiments. Towards this goal, we propose FineBio, a new fine-grained video dataset of people performing biological experiments. The dataset consists of multi-view videos of 32 participants performing mock biological experiments with a total duration of 14.5 hours. One experiment forms a hierarchical structure, where a protocol consists of several steps, each further decomposed into a set of atomic operations. The uniqueness of biological experiments is that while they require strict adherence to steps described in each protocol, there is freedom in the order of atomic operations. We provide hierarchical annotation on protocols, steps, atomic operations, object locations, and their manipulation states, providing new challenges for structured activity understanding and hand-object interaction recognition. To find out challenges on activity understanding in biological experiments, we introduce baseline models and results on four different tasks, including (i) step segmentation, (ii) atomic operation detection (iii) object detection, and (iv) manipulated/affected object detection. Dataset and code are available from https://github.com/aistairc/FineBio .
Takuma Yagi, Misaki Ohashi, Yifei Huang 0002, Ryosuke Furuta, Shungo Adachi, Totai Mitsuyama, Yoichi Sato 0001
Int. J. Comput. Vis.7
2025 Audio-visual localization based on spatial relative sound order
abstract
Abstract Sound localization is one of the essential tasks in audio-visual learning. Especially, stereo sound localization methods have been proposed to handle multiple sound sources. However, existing stereo-sound localization methods treat sound source localization as a segmentation task and, as a result, require costly annotation of segmentation masks. Another serious problem of the existing stereo-sound localization methods is that they have been trained and evaluated only in a controlled environment, such as a fixed camera and microphone setting with limited variability of scenes. Therefore, their performance on videos recorded in uncontrolled environments, such as in-the-wild videos from the Internet, has not been fully investigated. To address these problems, we propose a weakly supervised method as an extension of a typical stereo-sound localization method by utilizing the spatial relative order of sound sources in recorded videos. The proposed method solves the annotation problem by training the localization model using only sound category labels. Furthermore, our method utilizes the spatial relative order of the sound sources, which is not affected by specific recording settings, and thus can be effectively used for videos recorded in uncontrolled environments. We also collect stereo-recorded videos from YouTube to construct a new dataset to demonstrate the applicability of the proposed method to stereo sounds recorded in various environments. Our method enhances the localization performance by inserting a novel training step that exploits the relative order of sound sources into a typical audio-visual localization method in both existing and newly introduced audio-visual datasets.
Tomoya Sato, Yusuke Sugano, Yoichi Sato 0001
Mach. Vis. Appl.3
2025 Editorial Introduction to the ICCV 2021 Special Section
Dima Damen, Tal Hassner, Christopher Joseph Pal, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Ego4D: Around the World in 3,600 Hours of Egocentric Video
abstract
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception.
Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Devansh Kukreja, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik
IEEE Trans. Pattern Anal. Mach. Intell.80
2025 Prompt-Augmented Boundary Attentive Learning for Weakly Supervised Temporal Sentence Grounding
abstract
Weakly supervised temporal sentence grounding aims to temporally locate events described by a sentence in a video, relying solely on video-level visual-language correspondences. Because of the absence of precise boundary information, existing works primarily focus on multiple instance learning methods to establish segment-level video-language alignment. In this work, we propose Prompt-augmented Boundary Attentive Learning (PBAL) to enable the explicit modeling of the segment boundaries in a weakly supervised context. To represent the boundaries with sentences, we first generate sentences describing the start and end of an event, leveraging the capabilities of large language models (LLMs). With the augmented sentences, we then model the boundary-level video-language correspondence using a novel boundary-attentive learning module. This module generates probability maps of the starting and ending points, and is learned through boundary type prediction and self-supervised reconstruction. Experiments on two standard datasets, Charades-STA [1] and ActivityNet Captions [2] demonstrate PBAL’s state-of-the-art performance. The results of our ablation study further demonstrate the effectiveness of our boundary-attentive learning and prompt augmentation techniques.
Zhehao Zhu, Yifei Huang 0002, Mingfang Zhang 0002, Liangyang Ouyang, Yoichi Sato 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
abstract
We present Ego-Exo4D, a diverse, large-scale multi-modal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured ego-centric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is un-precedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions-including a novel “expert commentary” done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community.
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Makoto Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zachary Chavis, Joya Chen, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong 0002, María Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang 0020, Md Mohaiminul Islam, Suyog Dutt Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J. Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh K. Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina González, Prince Gupta, Jiabo Hu, Yifei Huang 0002, Yiming Huang 0011, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu 0007, Mi Luo, Zhengyi Luo 0002, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Kiran K. Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao 0002, Ziwei Zhao 0003, Zhifan Zhu 0001, Jeff Zhuo, Pablo Andrés Arbeláez, Gedas Bertasius, Dima Damen, Jakob J. Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard A. Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Manolis Savva, Jianbo Shi, Mike Zheng Shout, Michael Wray
CVPR96
2024 Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation
abstract
The pursuit of accurate 3D hand pose estimation stands as a keystone for understanding human activity in the realm of egocentric vision. The majority of existing estimation methods still rely on single-view images as input, leading to potential limitations, e.g., limited field-of-view and ambiguity in depth. To address these problems, adding another camera to better capture the shape of hands is a practi-cal direction. However, existing multi-view hand pose estimation methods suffer from two main drawbacks: 1) Re-quiring multi-view annotations for training, which are ex-pensive. 2) During testing, the model becomes inapplica-ble if camera parameters/layout are not the same as those used in training. In this paper, we propose a novel Single-to-Dual-view adaptation (S2DHand) solution that adapts a pretrained single-view estimator to dual views. Compared with existing multi-view training methods, 1) our adaptation process is unsupervised, eliminating the need for multi-view annotation. 2) Moreover, our method can handle arbitrary dual-view pairs with unknown camera parameters, making the model applicable to diverse camera settings. Specifically, S2DHand is built on certain stereo constraints, including pairwise cross-view consensus and invariance of transformation between both views. These two stereo constraints are used in a complementary manner to gen-erate pseudo-labels, allowing reliable adaptation. Evalu-ation results reveal that S2DHand achieves significant improvements on arbitrary camera pairs under both in-dataset and cross-dataset settings, and outperforms existing adaptation methods with leading performance. Project page: https://github.com/ut-vision/S2DHand.
Ruicong Liu, Takehiko Ohkawa, Mingfang Zhang 0002, Yoichi Sato 0001
CVPR4
2024 Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
Zicong Fan, Takehiko Ohkawa, Linlin Yang 0001, Nie Lin, Zhishan Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Feng Lu 0005, Karim Abou Zeid, Bastian Leibe, Jeongwan On, Seungryul Baek, Saurabh Gupta 0001, Yoichi Sato 0001, Otmar Hilliges, Hyung Jin Chang, Angela Yao
ECCV (25)21
2024 WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-Grained Spatial-Temporal Understanding
Quan Kong, Yuki Kawana, Rajat Saini, Jingjing Pan, Ta Gu, Yohei Ozao, Istvan Balazs Opra, Yoichi Sato 0001, Norimasa Kobori
ECCV (76)9
2024 ActionVOS: Actions as Prompts for Video Object Segmentation
Liangyang Ouyang, Ruicong Liu, Yifei Huang 0002, Ryosuke Furuta, Yoichi Sato 0001
ECCV (10)5
2024 Masked Video and Body-Worn IMU Autoencoder for Egocentric Action Recognition
Mingfang Zhang 0002, Yifei Huang 0002, Ruicong Liu, Yoichi Sato 0001
ECCV (18)4
2024 Matching Compound Prototypes for Few-Shot Action Recognition
abstract
Abstract The task of few-shot action recognition aims to recognize novel action classes using only a small number of labeled training samples. How to better describe the action in each video and how to compare the similarity between videos are two of the most critical factors in this task. Directly describing the video globally or by its individual frames cannot well represent the spatiotemporal dependencies within an action. On the other hand, naively matching the global representations of two videos is also not optimal since action can happen at different locations in a video with different speeds. In this work, we propose a novel approach that describes each video using multiple types of prototypes and then computes the video similarity with a particular matching strategy for each type of prototypes. To better model the spatiotemporal dependency, we describe the video by generating prototypes that model the multi-level spatiotemporal relations via transformers. There are a total of three types of prototypes. The first type of prototypes are trained to describe specific aspects of the action in the video e.g., the start of the action, regardless of its timestamp. These prototypes are directly matched one-to-one between two videos to compare their similarity. The second type of prototypes are the timestamp-centered prototypes that are trained to focus on specific timestamps of the video. To deal with the temporal variation of actions in a video, we apply bipartite matching to allow the matching of prototypes of different timestamps. The third type of prototypes are generated from the timestamp-centered prototypes, which regularize their temporal consistency while serving as an auxiliary summarization of the whole video. Experiments demonstrate that our proposed method achieves state-of-the-art results on multiple benchmarks.
Yifei Huang 0002, Lijin Yang, Guo Chen 0006, Hongjie Zhang 0002, Feng Lu 0005, Yoichi Sato 0001
Int. J. Comput. Vis.6
2024 Simultaneous control of head pose and expressions in 3D facial keypoint-based GAN
abstract
Abstract In this work, we present a novel method for simultaneously controlling the head pose and the facial expressions of a given input image using a 3D keypoint-based GAN. Existing methods for controlling head pose and expressions simultaneously are not suitable for real images, or they generate unnatural results because it is not trivial to capture head pose (large changes) and expressions (small changes) simultaneously. In this work, we achieve simultaneous control of head pose and facial expressions by introducing 3D facial keypoints for GAN-based facial image synthesis, unlike the existing 2D landmark-based approach. As a result, our method can handle both large variations due to different head poses and subtle variations due to changing facial expressions faithfully. Furthermore, our model takes audio input as an additional modality for further enhancing the quality of generated images. Our model was evaluated on the VoxCeleb2 dataset to demonstrate its state-of-the-art performance for both facial reenactment and facial image manipulation tasks, and our model tends not to be affected by the driving images.
Tomoyuki Hatakeyama, Ryosuke Furuta, Yoichi Sato 0001
Multim. Tools Appl.3
2024 Semantic Image Segmentation by Dynamic Discriminative Prototypes
abstract
Semantic segmentation achieves significant success through large-scale training data. Meanwhile, few-shot semantic segmentation was proposed to segment image regions of novel classes through few labeled testing data. However, it ignores classes previously learned from training data. This paper proposes a new segmentation framework called segmentation by dynamic prototype (SDP), which can simultaneously segment the image regions of base classes learned from many training data and novel classes learned from a few testing data. SDP performs segmentation by searching for the nearest prototype of each pixel's features. Different prototypes are representative features for different classes. In testing, SDP dynamically constructs novel prototypes according to support images while maintaining base prototypes learned from training images. The main challenge of SDP is how to achieve intra-image compactness, intra-class compactness, and inter-class separability simultaneously in the feature space. To tackle these challenges, we first introduce a discriminative pixelwise feature and prototype training method to improve the above three types of feature discriminability. We then introduce a mask refinement process in testing, which refines support images' masks to extract more separable novel prototypes. In addition, we introduce a prototype adaptation process in testing, which allows all prototypes to adapt to query images to reduce the prototypes' intra-class variances. Our approach achieves state-of-the-art results for novel class segmentation against existing few-shot semantic segmentation methods on PASCAL-5iand COCO-20ibenchmarks in combination with superior runtime efficiency. In addition, our method can maintain strong results in the segmentation of base classes.
Kaipeng Zhang, Yoichi Sato 0001
IEEE Trans. Multim.2
2023 Proposal-based Temporal Action Localization with Point-level Supervision
Yifei Huang 0002, Ryosuke Furuta, Yoichi Sato 0001
BMVC4
2023 Structural Multiplane Image: Bridging Neural View Synthesis and 3D Reconstruction
abstract
The Multiplane Image (MPI), containing a set of fronto-parallel$RGB_{\alpha}$layers, is an effective and efficient representation for view synthesis from sparse inputs. Yet, its fixed structure limits the performance, especially for surfaces imaged at oblique angles. We introduce the Structural MPI (S-MPI), where the plane structure approximates 3D scenes concisely. Conveying$RGB_{\alpha}$contexts with geometrically-faithful structures, the S-MPI directly bridges view synthe-sis and 3D reconstruction. It can not only overcome the critical limitations of MPI, i.e., discretization artifacts from sloped surfaces and abuse of redundant layers, and can also acquire planar 3D reconstruction. Despite the intu-ition and demand of applying S-MPI, great challenges are introduced,$e.g$., high-fidelity approximation for both$RGB_{\alpha}$layers and plane poses, multi-view consistency, non-planar regions modeling, and efficient rendering with intersected planes. Accordingly, we propose a transformer-based network based on a segmentation model [4]. It predicts compact and expressive S-MPI layers with their corresponding masks, poses, and$RGB_{\alpha}$contexts. Non-planar regions are inclusively handled as a special case in our unified frame-work. Multi-view consistency is ensured by sharing global proxy embeddings, which encode plane-level features cov-ering the complete 3D scenes with aligned coordinates. In-tensive experiments show that our method outperforms both previous state-of-the-art MPI-based view synthesis methods and planar reconstruction methods.
Mingfang Zhang 0002, Jinglu Wang, Xiao Li 0030, Yifei Huang 0002, Yoichi Sato 0001, Yan Lu 0001
CVPR5
2023 Weakly Supervised Temporal Sentence Grounding with Uncertainty-Guided Self-training
abstract
The task of weakly supervised temporal sentence grounding aims at finding the corresponding temporal moments of a language description in the video, given video-language correspondence only at video-level. Most existing works select mismatched video-language pairs as negative samples and train the model to generate better positive proposals that are distinct from the negative ones. However, due to the complex temporal structure of videos, proposals distinct from the negative ones may correspond to several video segments but not necessarily the correct ground truth. To alleviate this problem, we propose an uncertainty-guided self-training technique to provide extra self-supervision signal to guide the weakly-supervised learning. The self-training process is based on teacher-student mutual learning with weak-strong augmentation, which enables the teacher network to generate relatively more reliable outputs compared to the student network, so that the student network can learn from the teacher's output. Since directly applying existing self-training methods in this task easily causes error accumulation, we specifically design two techniques in our selftraining method: (1) we construct a Bayesian teacher network, leveraging its uncertainty as a weight to suppress the noisy teacher supervisory signals; (2) we leverage the cycle consistency brought by temporal data augmentation to perform mutual learning between the two networks. Experiments demonstrate our method's superiority on Charades-STA and ActivityNet Captions datasets. We also show in the experiment that our self-training method can be applied to improve the performance of multiple backbone methods.
Yifei Huang 0002, Lijin Yang, Yoichi Sato 0001
CVPR3
2023 DeCo: Decomposition and Reconstruction for Compositional Temporal Grounding via Coarse-to-Fine Contrastive Ranking
abstract
Understanding dense action in videos is a fundamental challenge towards the generalization of vision models. Several works show that compositionality is key to achieving generalization by combining known primitive elements, especially for handling novel composited structures. Compositional temporal grounding is the task of localizing dense action by using known words combined in novel ways in the form of novel query sentences for the actual grounding. In recent works, composition is assumed to be learned from pairs of whole videos and language embeddings through large scale self-supervised pre-training. Alternatively, one can process the video and language into word-level primitive elements, and then only learn fine-grained semantic correspondences. Both approaches do not consider the granularity of the compositions, where different query granularity corresponds to different video segments. Therefore, a good compositional representation should be sensitive to different video and query granularity. We propose a method to learn a coarse-to-fine compositional representation by decomposing the original query sentence into different granular levels, and then learning the correct correspondences between the video and recombined queries through a contrastive ranking constraint. Additionally, we run temporal boundary prediction in a coarse-to-fine manner for precise grounding boundary detection. Experiments are performed on two datasets, Charades-CG and ActivityNet-CG, showing the superior compositional generalizability of our approach.
Lijin Yang, Quan Kong, Hsuan-Kung Yang, Wadim Kehl, Yoichi Sato 0001, Norimasa Kobori
CVPR5
2023 Image Cropping under Design Constraints
abstract
Image cropping is essential in image editing for obtaining a compositionally enhanced image. In display media, image cropping is a prospective technique for automatically creating media content. However, image cropping for media contents is often required to satisfy various constraints, such as an aspect ratio and blank regions for placing texts or objects. We call this problem image cropping under design constraints. To achieve image cropping under design constraints, we propose a score function-based approach, which computes scores for cropped results whether aesthetically plausible and satisfies design constraints. We explore two derived approaches, a proposal-based approach, and a heatmap-based approach, and we construct a dataset for evaluating the performance of the proposed approaches on image cropping under design constraints. In experiments, we demonstrate that the proposed approaches outperform a baseline, and we observe that the proposal-based approach is better than the heatmap-based approach under the same computation cost, but the heatmap-based approach leads to better scores by increasing computation cost. The experimental results indicate that balancing aesthetically plausible regions and satisfying design constraints is not a trivial problem and requires sensitive balance, and both proposed approaches are reasonable alternatives.
Takumi Nishiyasu, Wataru Shimoda, Yoichi Sato 0001
MMAsia3
2023 Fine-grained Affordance Annotation for Egocentric Hand-Object Interaction Videos
abstract
Object affordance is an important concept in hand-object interaction, providing information on action possibilities based on human motor capacity and objects’ physical property thus benefiting tasks such as action anticipation and robot imitation learning. However, the definition of affordance in existing datasets often: 1) mix up affordance with object functionality; 2) confuse affordance with goal-related action; and 3) ignore human motor capacity. This paper proposes an efficient annotation scheme to address these issues by combining goal-irrelevant motor actions and grasp types as affordance labels and introducing the concept of mechanical action to represent the action possibilities between two objects. We provide new annotations by applying this scheme to the EPIC-KITCHENS dataset and test our annotation with tasks such as affordance recognition, hand-object interaction hotspots prediction, and cross-domain evaluation of affordance. The results show that models trained with our annotation can distinguish affordance from other concepts, predict fine-grained interaction possibilities on objects, and generalize through different do-mains.
Zecheng Yu, Yifei Huang 0002, Ryosuke Furuta, Takuma Yagi, Yusuke Goutsu, Yoichi Sato 0001
WACV6
2023 Efficient Annotation and Learning for 3D Hand Pose Estimation: A Survey
abstract
Abstract In this survey, we present a systematic review of 3D hand pose estimation from the perspective of efficient annotation and learning. 3D hand pose estimation has been an important research area owing to its potential to enable various applications, such as video understanding, AR/VR, and robotics. However, the performance of models is tied to the quality and quantity of annotated 3D hand poses. Under the status quo, acquiring such annotated 3D hand poses is challenging, e.g., due to the difficulty of 3D annotation and the presence of occlusion. To reveal this problem, we review the pros and cons of existing annotation methods classified as manual, synthetic-model-based, hand-sensor-based, and computational approaches. Additionally, we examine methods for learning 3D hand poses when annotated data are scarce, including self-supervised pretraining, semi-supervised learning, and domain adaptation. Based on the study of efficient annotation and learning, we further discuss limitations and possible future directions in this field.
Takehiko Ohkawa, Ryosuke Furuta, Yoichi Sato 0001
Int. J. Comput. Vis.3
2022 Ego4D: Around the World in 3, 000 Hours of Egocentric Video
abstract
We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik
CVPR79
2022 Interact before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action Recognition
abstract
Unsupervised domain adaptive video action recognition aims to recognize actions of a target domain using a model trained with only out-of-domain (source) annotations. The inherent complexity of videos makes this task challenging but also provides ground for leveraging multi-modal inputs (e.g., RGB, Flow, Audio). Most previous works utilize the multi-modal information by either aligning each modality individually or learning representation via cross-modal self-supervision. Different from previous works, we find that the cross-domain alignment can be more effectively done by using cross-modal interaction first. Cross-modal knowledge interaction allows other modalities to supplement missing transferable information because of the cross-modal complementarity. Also, the most transferable aspects of data can be highlighted using cross-modal consensus. In this work, we present a novel model that jointly considers these two characteristics for domain adaptive action recognition. We achieve this by implementing two modules, where the first module exchanges complementary transferable information across modalities through the semantic space, and the second module finds the most transferable spatial region based on the consensus of all modalities. Extensive experiments validate that our proposed method can significantly outperform state-of-the-art methods on multiple benchmark datasets, including the complex fine-grained dataset EPIC-Kitchens-100.
Lijin Yang, Yifei Huang 0002, Yusuke Sugano, Yoichi Sato 0001
CVPR4
2022 Compound Prototype Matching for Few-Shot Action Recognition
Yifei Huang 0002, Lijin Yang, Yoichi Sato 0001
ECCV (4)3
2022 CompNVS: Novel View Synthesis with Scene Completion
Zuoyue Li, Tianxing Fan, Zhenqiang Li 0002, Zhaopeng Cui, Yoichi Sato 0001, Marc Pollefeys, Martin R. Oswald
ECCV (1)5
2022 Domain Adaptive Hand Keypoint and Pixel Localization in the Wild
Takehiko Ohkawa, Yu-Jhe Li, Qichen Fu, Ryosuke Furuta, Kris Makoto Kitani, Yoichi Sato 0001
ECCV (9)6
2022 Surgical Skill Assessment via Video Semantic Aggregation
Zhenqiang Li 0002, Lin Gu 0003, Weimin Wang 0007, Ryosuke Nakamura, Yoichi Sato 0001
MICCAI (8)5
2022 Spatio-Temporal Perturbations for Video Attribution
abstract
The attribution method provides a direction for interpreting opaque neural networks in a visual way by identifying and visualizing the input regions/pixels that dominate the output of a network. Regarding the attribution method for visually explaining video understanding networks, it is challenging because of the unique spatiotemporal dependencies existing in video inputs and the special 3D convolutional or recurrent structures of video understanding networks. However, most existing attribution methods focus on explaining networks taking a single image as input and a few works specifically devised for video attribution come short of dealing with diversified structures of video understanding networks. In this paper, we investigate a generic perturbation-based attribution method that is compatible with diversified video understanding networks. Besides, we propose a novel regularization term to enhance the method by constraining the smoothness of its attribution results in both spatial and temporal dimensions. In order to assess the effectiveness of different video attribution methods without relying on manual judgement, we introduce reliable objective metrics which are checked by a newly proposed reliability measurement. We verified the effectiveness of our method by both subjective and objective evaluation and comparison with multiple significant attribution methods.
Zhenqiang Li 0002, Weimin Wang 0007, Zuoyue Li, Yifei Huang 0002, Yoichi Sato 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 GO-Finder: A Registration-free Wearable System for Assisting Users in Finding Lost Hand-held Objects
abstract
People spend an enormous amount of time and effort looking for lost objects. To help remind people of the location of lost objects, various computational systems that provide information on their locations have been developed. However, prior systems for assisting people in finding objects require users to register the target objects in advance. This requirement imposes a cumbersome burden on the users, and the system cannot help remind them of unexpectedly lost objects. We propose GO-Finder (“Generic Object Finder”), a registration-free wearable camera-based system for assisting people in finding an arbitrary number of objects based on two key features: automatic discovery of hand-held objects and image-based candidate selection. Given a video taken from a wearable camera, GO-Finder automatically detects and groups hand-held objects to form a visual timeline of the objects. Users can retrieve the last appearance of the object by browsing the timeline through a smartphone app. We conducted user studies to investigate how users benefit from using GO-Finder. In the first study, we asked participants to perform an object retrieval task and confirmed improved accuracy and reduced mental load in the object search task by providing clear visual cues on object locations. In the second study, the system’s usability on a longer and more realistic scenario was verified, accompanied by an additional feature of context-based candidate filtering. Participant feedback suggested the usefulness of GO-Finder also in realistic scenarios where more than one hundred objects appear.
Takuma Yagi, Takumi Nishiyasu, Kunimasa Kawasaki, Moe Matsuki, Yoichi Sato 0001
ACM Trans. Interact. Intell. Syst.5
2021 Leveraging Human Selective Attention for Medical Image Analysis with Limited Training Data
Yifei Huang 0002, Lijin Yang, Lin Gu 0003, Yingying Zhu 0004, Hirofumi Seo, Qiuming Meng, Tatsuya Harada, Yoichi Sato 0001
BMVC9
2021 Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction
Takuma Yagi, Md Tasnimul Hasan, Yoichi Sato 0001
BMVC3
2021 Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips
Lijin Yang, Yifei Huang 0002, Yusuke Sugano, Yoichi Sato 0001
BMVC4
2021 Unsupervised Common Particular Object Discovery and Localization by Analyzing a Match Graph
abstract
Although the unsupervised discovery and localization of common objects from within a set of images has received considerable attention, the difficulty of this task means that current methods are not sufficiently accurate. This paper describes an unsupervised method that more accurately discovers and localizes common particular objects within a set of images. First, our method constructs a match graph that precisely represents the relationship among particular object proposals extracted from images. The match graph necessarily contains communities, i.e., subgraphs where object proposals with identical subjects are densely connected by edges, indicating matches between local image features in the proposals. Our method determines the common particular objects by scoring the object proposals using a novel node centrality measure for graphs with a community structure, which we call the restrained random-walk centrality. Experiments demonstrate that our method is considerably more accurate than previous approaches on a particular object image set.
Makoto Okuda, Shin'ichi Satoh 0001, Yoichi Sato 0001, Yutaka Kidawara
ICASSP3
2021 GO-Finder: A Registration-Free Wearable System for Assisting Users in Finding Lost Objects via Hand-Held Object Discovery
abstract
People spend an enormous amount of time and effort looking for lost objects. To help remind people of the location of lost objects, various computational systems that provide information on their locations have been developed. However, prior systems for assisting people in finding objects require users to register the target objects in advance. This requirement imposes a cumbersome burden on the users, and the system cannot help remind them of unexpectedly lost objects. We propose GO-Finder (“Generic Object Finder”), a registration-free wearable camera based system for assisting people in finding an arbitrary number of objects based on two key features: automatic discovery of hand-held objects and image-based candidate selection. Given a video taken from a wearable camera, Go-Finder automatically detects and groups hand-held objects to form a visual timeline of the objects. Users can retrieve the last appearance of the object by browsing the timeline through a smartphone app. We conducted a user study to investigate how users benefit from using GO-Finder and confirmed improved accuracy and reduced mental load regarding the object search task by providing clear visual cues on object locations.
Takuma Yagi, Takumi Nishiyasu, Kunimasa Kawasaki, Moe Matsuki, Yoichi Sato 0001
IUI5
2021 Neural Routing by Memory
abstract
Recent Convolutional Neural Networks (CNNs) have achieved significant success by stacking multiple convolutional blocks, named procedures in this paper, to extract semantic features. However, they use the same procedure sequence for all inputs, regardless of the intermediate features.This paper proffers a simple yet effective idea of constructing parallel procedures and assigning similar intermediate features to the same specialized procedures in a divide-and-conquer fashion. It relieves each procedure's learning difficulty and thus leads to superior performance. Specifically, we propose a routing-by-memory mechanism for existing CNN architectures. In each stage of the network, we introduce parallel Procedural Units (PUs). A PU consists of a memory head and a procedure. The memory head maintains a summary of a type of features. For an intermediate feature, we search its closest memory and forward it to the corresponding procedure in both training and testing. In this way, different procedures are tailored to different features and therefore tackle them better.Networks with the proposed mechanism can be trained efficiently using a four-step training strategy. Experimental results show that our method improves VGGNet, ResNet, and EfficientNet's accuracies on Tiny ImageNet, ImageNet, and CIFAR-100 benchmarks with a negligible extra computational cost.
Kaipeng Zhang, Zhenqiang Li 0002, Zhifeng Li 0001, Wei Liu 0005, Yoichi Sato 0001
NeurIPS5
2021 Towards Visually Explaining Video Understanding Networks with Perturbation
abstract
"Making black box models explainable " is a vital problem that accompanies the development of deep learning networks. For networks taking visual information as input, one basic but challenging explanation method is to identify and visualize the input pixels/regions that dominate the network's prediction. However, most existing works focus on explaining networks taking a single image as input and do not consider the temporal relationship that exists in videos. Providing an easy-to-use visual explanation method that is applicable to diversified structures of video understanding networks still remains an open challenge. In this paper, we investigate a generic perturbation-based method for visually explaining video understanding networks. Besides, we propose a novel loss function to enhance the method by constraining the smoothness of its results in both spatial and temporal dimensions. The method enables the comparison of explanation results between different network structures to become possible and can also avoid generating the pathological adversarial explanations for video inputs. Experimental comparison results verified the effectiveness of our method.
Zhenqiang Li 0002, Weimin Wang 0007, Zuoyue Li, Yifei Huang 0002, Yoichi Sato 0001
WACV5
2021 Community Detection Using Restrained Random-Walk Similarity
abstract
In this paper, we propose a restrained random-walk similarity method for detecting the community structures of graphs. The basic premise of our method is that the starting vertices of finite-length random walks are judged to be in the same community if the walkers pass similar sets of vertices. This idea is based on our consideration that a random walker tends to move in the community including the walker's starting vertex for some time after starting the walk. Therefore, the sets of vertices passed by random walkers starting from vertices in the same community must be similar. The idea is reinforced with two conditions. First, we exclude abnormal random walks. Random walks that depart from each vertex are executed many times, and vertices that are rarely passed by the walkers are excluded from the set of vertices that the walkers may pass. Second, we forcibly restrain random walks to an appropriate length. In our method, a random walk is terminated when the walker repeatedly visits vertices that they have already passed. Experiments on real-world networks demonstrate that our method outperforms previous techniques in terms of accuracy.
Makoto Okuda, Shin'ichi Satoh 0001, Yoichi Sato 0001, Yutaka Kidawara
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Support Strategies for Remote Guides in Assisting People with Visual Impairments for Effective Indoor Navigation
abstract
People with visual impairments often require mobility assistance of sighted guides but they are not always available. Recent technological strides have opened up new directions for sighted guidance services, assigning guides from a network of remote workers to provide real-time assistance via audio/video communication. However, little has been known regarding desirable support characteristics of remote guides or challenges experienced in guide practices without the requisite expertise. To recommend support strategies that contribute to facilitating a successful platform for remote sighted guidance, this paper presents a comparative study of the performance of trained and untrained sighted guides who are recruited for a remote scenario in assisting people with visual impairments in indoor navigation. As an outcome of this research, we provide a deeper understanding of design opportunities for HCI to scaffold requirements of remote guides, such that their collaborative efforts and environmental knowledge influence the user experience. Based on our empirical insights, we suggest to develop the expertise of remote guides through: a) preliminary guidance cooperation awareness b) guidelines for verbal description methods, and c) approaches to compensate for the lack of environmental knowledge.
Rie Kamikubo, Naoya Kato, Keita Higuchi, Ryo Yonetani, Yoichi Sato 0001
CHI5
2020 Generalizing Hand Segmentation in Egocentric Videos With Uncertainty-Guided Model Adaptation
abstract
Although the performance of hand segmentation in egocentric videos has been significantly improved by using CNNs, it still remains a challenging issue to generalize the trained models to new domains, e.g., unseen environments. In this work, we solve the hand segmentation generalization problem without requiring segmentation labels in the target domain. To this end, we propose a Bayesian CNN-based model adaptation framework for hand segmentation, which introduces and considers two key factors: 1) prediction uncertainty when the model is applied in a new domain and 2) common information about hand shapes shared across domains. Consequently, we propose an iterative self-training method for hand segmentation in the new domain, which is guided by the model uncertainty estimated by a Bayesian CNN. We further use an adversarial component in our framework to utilize shared information about hand shapes to constrain the model adaptation process. Experiments on multiple egocentric datasets show that the proposed method significantly improves the generalization performance of hand segmentation.
Minjie Cai, Feng Lu 0005, Yoichi Sato 0001
CVPR3
2020 Improving Action Segmentation via Graph-Based Temporal Reasoning
abstract
Temporal relations among multiple action segments play an important role in action segmentation especially when observations are limited (e.g., actions are occluded by other objects or happen outside a field of view). In this paper, we propose a network module called Graph-based Temporal Reasoning Module (GTRM) that can be built on top of existing action segmentation models to learn the relation of multiple action segments in various time spans. We model the relations by using two Graph Convolution Networks (GCNs) where each node represents an action segment. The two graphs have different edge properties to account for boundary regression and classification tasks, respectively. By applying graph convolution, we can update each node's representation based on its relation with neighboring nodes. The updated representation is then used for improved action segmentation. We evaluate our model on the challenging egocentric datasets namely EGTEA and EPIC-Kitchens, where actions may be partially observed due to the viewpoint restriction. The results show that our proposed GTRM outperforms state-of-the-art action segmentation models by a large margin. We also demonstrate the effectiveness of our model on two third-person video datasets, the 50Salads dataset and the Breakfast dataset.
Yifei Huang 0002, Yusuke Sugano, Yoichi Sato 0001
CVPR3
2020 Hierarchical Gaussian Descriptors with Application to Person Re-Identification
abstract
Describing the color and textural information of a person image is one of the most crucial aspects of person re-identification (re-id). Although a covariance descriptor has been successfully applied to person re-id, it loses the local structure of a region and mean information of pixel features, both of which tend to be the major discriminative information for person re-id. In this paper, we present novel meta-descriptors based on a hierarchical Gaussian distribution of pixel features, in which both mean and covariance information are included in patch and region level descriptions. More specifically, the region is modeled as a set of multiple Gaussian distributions, each of which represents the appearance of a local patch. The characteristics of the set of Gaussian distributions are again described by another Gaussian distribution. Because the space of Gaussian distribution is not a linear space, we embed the parameters of the distribution into a point of Symmetric Positive Definite (SPD) matrix manifold in both steps. We show, for the first time, that normalizing the scale of the SPD matrix enhances the hierarchical feature representation on this manifold. Additionally, we develop feature norm normalization methods with the ability to alleviate the biased trends that exist on the SPD matrix descriptors. The experimental results conducted on five public datasets indicate the effectiveness of the proposed descriptors and the two types of normalizations.
Tetsu Matsukawa, Takahiro Okabe, Einoshin Suzuki, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 An Ego-Vision System for Discovering Human Joint Attention
abstract
Joint attention often happens during social interactions, in which individuals share focus on the same object. This article proposes an egocentric vision-based system (ego-vision system) that aims to discover the objects looked at jointly by a group of persons engaged in interactive activities. The proposed system relies on a collection of wearable eye-tracking cameras that provide an egocentric view of the interaction scenes as well as points-of-gaze measurement of each participant. Technically in our system, we develop a hierarchical conditional random field (CRF) based graphical model that can temporally localize joint attention periods and spatially segment objects of joint attention. By solving these two coupled tasks together in an iterative optimization procedure, we show that human joint attention can be reliably discovered from videos even with cluttered background and noisy gaze measurement. A new dataset of joint attention is collected and annotated for evaluating the two tasks of joint attention where two to four persons are involved. Experimental results demonstrate that our approach achieves state-of-the-art performance on both tasks of spatial segmentation and temporal localization of joint attention.
Yifei Huang 0002, Minjie Cai, Yoichi Sato 0001
IEEE Trans. Hum. Mach. Syst.3
2020 Learning Context-dependent Personal Preferences for Adaptive Recommendation
abstract
We propose two online-learning algorithms for modeling the personal preferences of users of interactive systems. The proposed algorithms leverage user feedback to estimate user behavior and provide personalized adaptive recommendation for supporting context-dependent decision-making. We formulate preference modeling as online prediction algorithms over a set of learned policies, i.e., policies generated via supervised learning with interaction and context data collected from previous users. The algorithms then adapt to a target user by learning the policy that best predicts that user’s behavior and preferences. We also generalize the proposed algorithms for a more challenging learning case in which they are restricted to a limited number of trained policies at each timestep, i.e., for mobile settings with limited resources. While the proposed algorithms are kept general for use in a variety of domains, we developed an image-filter-selection application. We used this application to demonstrate how the proposed algorithms can quickly learn to match the current user’s selections. Based on these evaluations, we show that (1) the proposed algorithms exhibit better prediction accuracy compared to traditional supervised learning and bandit algorithms, (2) our algorithms are robust under challenging limited prediction settings in which a smaller number of expert policies is assumed. Finally, we conducted a user study to demonstrate how presenting users with the prediction results of our algorithms significantly improves the efficiency of the overall interaction experience.
Keita Higuchi, Hiroki Tsuchida, Eshed Ohn-Bar, Yoichi Sato 0001, Kris Makoto Kitani
ACM Trans. Interact. Intell. Syst.4
2020 Gaze Estimation by Exploring Two-Eye Asymmetry
abstract
Eye gaze estimation is increasingly demanded by recent intelligent systems to facilitate a range of interactive applications. Unfortunately, learning the highly complicated regression from a single eye image to the gaze direction is not trivial. Thus, the problem is yet to be solved efficiently. Inspired by the two-eye asymmetry as two eyes of the same person may appear uneven, we propose the face-based asymmetric regression-evaluation network (FARE-Net) to optimize the gaze estimation results by considering the difference between left and right eyes. The proposed method includes one face-based asymmetric regression network (FAR-Net) and one evaluation network (E-Net). The FAR-Net predicts 3D gaze directions for both eyes and is trained with the asymmetric mechanism, which asymmetrically weights and sums the loss generated by two-eye gaze directions. With the asymmetric mechanism, the FAR-Net utilizes the eyes that can achieve high performance to optimize network. The E-Net learns the reliabilities of two eyes to balance the learning of the asymmetric mechanism and symmetric mechanism. Our FARENet achieves leading performances on MPIIGaze, EyeDiap and RT-Gene datasets. Additionally, we investigate the effectiveness of FARE-Net by analyzing the distribution of errors and ablation study.
Yihua Cheng, Xucong Zhang, Feng Lu 0005, Yoichi Sato 0001
IEEE Trans. Image Process.4
2020 Mutual Context Network for Jointly Estimating Egocentric Gaze and Action
abstract
In this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context: the information from gaze prediction facilitates action recognition and vice versa. Our assumption is that during the procedure of performing a manipulation task, on the one hand, what a person is doing determines where the person is looking at. On the other hand, the gaze location reveals gaze regions which contain important and information about the undergoing action and also the non-gaze regions that include complimentary clues for differentiating some fine-grained actions. We propose a novel mutual context network (MCN) that jointly learns action-dependent gaze prediction and gaze-guided action recognition in an end-to-end manner. Experiments on multiple egocentric video datasets demonstrate that our MCN achieves state-of-the-art performance of both gaze prediction and action recognition. The experiments also show that action-dependent gaze patterns could be learned with our method.
Yifei Huang 0002, Minjie Cai, Zhenqiang Li 0002, Feng Lu 0005, Yoichi Sato 0001
IEEE Trans. Image Process.5
2019 BBeep: A Sonic Collision Avoidance System for Blind Travellers and Nearby Pedestrians
abstract
We present an assistive suitcase system, BBeep, for supporting blind people when walking through crowded environments. BBeep uses pre-emptive sound notifications to help clear a path by alerting both the user and nearby pedestrians about the potential risk of collision. BBeep triggers notifications by tracking pedestrians, predicting their future position in real-time, and provides sound notifications only when it anticipates a future collision. We investigate how different types and timings of sound affect nearby pedestrian behavior. In our experiments, we found that sound emission timing has a significant impact on nearby pedestrian trajectories when compared to different sound types. Based on these findings, we performed a real-world user study at an international airport, where blind participants navigated with the suitcase in crowded areas. We observed that the proposed system significantly reduces the number of imminent collisions.
Seita Kayukawa, Keita Higuchi, João Guerreiro 0002, Shigeo Morishima, Yoichi Sato 0001, Kris Makoto Kitani, Chieko Asakawa
CHI5
2019 CoSummary: adaptive fast-forwarding for surgical videos by detecting collaborative scenes using hand regions and gaze positions
abstract
This paper presents CoSummary, an adaptive video fast-forwarding technique for browsing surgical videos recorded by wearable cameras. Current wearable technologies allow us to record complex surgical skills, however, an efficient browsing technique for these videos is not well established. In order to assist browsing surgical videos, our study focuses on adaptively changing playback speeds through the learning and detecting collaborative scenes based on surgeon hand placement and gaze information. Our evaluation shows that the proposed method is able to highlight important collaborative scenes and skip less important scenes during surgical procedures. We have also performed a subjective study with surgeons in order to have professional feedback. The results confirmed the effectiveness of the proposed method in comparison to uniform video fast-forwarding.
Irshad Abibouraguimane, Kakeru Hagihara, Keita Higuchi, Yuta Itoh 0001, Yoichi Sato 0001, Tetsu Hayashida, Maki Sugimoto
IUI5
2019 Assisting group activity analysis through hand detection and identification in multiple egocentric videos
abstract
Research in group activity analysis has put attention to monitor the work and evaluate group and individual performance, which can be reflected towards potential improvements in future group interactions. As a new means to examine individual or joint actions in the group activity, our work investigates the potential of detecting and disambiguating hands of each person in first-person points-of-view videos. Based on the recent developments in automated hand-region extraction from videos, we develop a new multiple-egocentric-video browsing interface that gives easy access to the frames of 1) individual action when only the hands of the viewer are detected, 2) joint action when collective hands are detected, and 3) the viewer checking the others' action as only their hands are detected. We take the evaluation process to explore the effectiveness of our interface with proposed hand-related features which can help perceive actions of interests in the complex analysis of videos involving co-occurred behaviors of multiple people.
Nathawan Charoenkulvanich, Rie Kamikubo, Ryo Yonetani, Yoichi Sato 0001
IUI4
2018 Future Person Localization in First-Person Videos
abstract
We present a new task that predicts future locations of people observed in first-person videos. Consider a first-person video stream continuously recorded by a wearable camera. Given a short clip of a person that is extracted from the complete stream, we aim to predict that person's location in future frames. To facilitate this future person localization ability, we make the following three key observations: (a) First-person videos typically involve significant ego-motion which greatly affects the location of the target person in future frames; (b) Scales of the target person act as a salient cue to estimate a perspective effect in first-person videos; (c) First-person videos often capture people up-close, making it easier to leverage target poses (e.g., where they look) for predicting their future locations. We incorporate these three observations into a prediction framework with a multi-stream convolution-deconvolution architecture. Experimental results reveal our method to be effective on our new dataset as well as on a public social interaction dataset.
Takuma Yagi, Karttikeya Mangalam, Ryo Yonetani, Yoichi Sato 0001
CVPR4
2018 Predicting Gaze in Egocentric Video by Learning Task-Dependent Attention Transition
Yifei Huang 0002, Minjie Cai, Zhenqiang Li 0002, Yoichi Sato 0001
ECCV (4)4
2018 Visualizing Gaze Direction to Support Video Coding of Social Attention for Children with Autism Spectrum Disorder
abstract
This paper presents a novel interface to support video coding of social attention in the assessment of children with autism spectrum disorder. Video-based evaluations of social attention during therapeutic activities allow observers to find target behaviors while handling the ambiguity of attention. Despite the recent advances in computer vision-based gaze estimation methods, fully automatic recognition of social attention under diverse environments is still challenging. The goal of this work is to investigate an approach that uses automatic video analysis in a supportive manner for guiding human judgment. The proposed interface displays visualization of gaze estimation results on videos and provides GUI support to allow users to facilitate agreement between observers by defining social attention labels on the video timeline. Through user studies and expert reviews, we show how the interface helps observers perform video coding of social attention and how human judgment compensates for technical limitations of the automatic gaze analysis.
Keita Higuchi, Soichiro Matsuda, Rie Kamikubo, Takuya Enomoto, Yusuke Sugano, Junichi Yamamoto, Yoichi Sato 0001
IUI7
2018 Browsing Group First-Person Videos with 3D Visualization
abstract
This work presents a novel user interface applying 3D visualization to understand complex group activities from multiple first-person videos. The proposed interface is designed to assist video viewers to easily understand the collaborative relationships of group activity based on where the individual worker is located in a workspace and how multiple workers are positioned to one another during the group activity. More specifically, the interface not only shows all recorded first-person videos but also visualizes the 3D position and orientation of each view point (i.e., the 3D position of each worker wearing a head-mounted camera) with a reconstructed 3D model of the workspace. Our user study confirms that the 3D visualization helps video viewers to understand geometric information of a worker and collaborative relationships of group activity easily and accurately.
Yuki Sugita, Keita Higuchi, Ryo Yonetani, Rie Kamikubo, Yoichi Sato 0001
ISS5
2018 Editorial for ACCV'16 award papers
Shang-Hong Lai, Vincent Lepetit, Ko Nishino, Yoichi Sato 0001
Comput. Vis. Image Underst.4
2018 SymPS: BRDF Symmetry Guided Photometric Stereo for Shape and Light Source Estimation
abstract
We propose uncalibrated photometric stereo methods that address the problem due to unknown isotropic reflectance. At the core of our methods is the notion of "constrained half-vector symmetry" for general isotropic BRDFs. We show that such symmetry can be observed in various real-world materials, and it leads to new techniques for shape and light source estimation. Based on the 1D and 2D representations of the symmetry, we propose two methods for surface normal estimation; one focuses on accurate elevation angle recovery for surface normals when the light sources only cover the visible hemisphere, and the other for comprehensive surface normal optimization in the case that the light sources are also non-uniformly distributed. The proposed robust light source estimation method also plays an essential role to let our methods work in an uncalibrated manner with good accuracy. Quantitative evaluations are conducted with both synthetic and real-world scenes, which produce the state-of-the-art accuracy for all of the non-Lambertian materials in MERL database and the real-world datasets.
Feng Lu 0005, Xiaowu Chen 0001, Imari Sato, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Continuous 3D Label Stereo Matching Using Local Expansion Moves
abstract
We present an accurate stereo matching method using local expansion moves based on graph cuts. This new move-making scheme is used to efficiently infer per-pixel 3D plane labels on a pairwise Markov random field (MRF) that effectively combines recently proposed slanted patch matching and curvature regularization terms. The local expansion moves are presented as many -expansions defined for small grid regions. The local expansion moves extend traditional expansion moves by two ways: localization and spatial propagation. By localization, we use different candidate -labels according to the locations of local -expansions. By spatial propagation, we design our local -expansions to propagate currently assigned labels for nearby regions. With this localization and spatial propagation, our method can efficiently infer MRF models with a continuous label space using randomized search. Our method has several advantages over previous approaches that are based on fusion moves or belief propagation; it produces submodular moves deriving a subproblem optimality; it helps find good, smooth, piecewise linear disparity maps; it is suitable for parallelization; it can use cost-volume filtering techniques for accelerating the matching cost computations. Even using a simple pairwise MRF, our method is shown to have best performance in the Middlebury stereo benchmark V2 and V3.
Tatsunori Taniai, Yasuyuki Matsushita, Yoichi Sato 0001, Takeshi Naemura
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Ego-Surfing: Person Localization in First-Person Videos Using Ego-Motion Signatures
abstract
We envision a future time when wearable cameras are worn by the masses and recording first-person point-of-view videos of everyday life. While these cameras can enable new assistive technologies and novel research challenges, they also raise serious privacy concerns. For example, first-person videos passively recorded by wearable cameras will necessarily include anyone who comes into the view of a camera-with or without consent. Motivated by these benefits and risks, we developed a self-search technique tailored to first-person videos. The key observation of our work is that the egocentric head motion of a target person (i.e., the self) is observed both in the point-of-view video of the target and observer. The motion correlation between the target person's video and the observer's video can then be used to identify instances of the self uniquely. We incorporate this feature into the proposed approach that computes the motion correlation over densely-sampled trajectories to search for a target individual in observer videos. Our approach significantly improves self-search performance over several well-known face detectors and recognizers. Furthermore, we show how our approach can enable several practical applications such as privacy filtering, target video retrieval, and social group clustering.
Ryo Yonetani, Kris Makoto Kitani, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Hyperspectral Image Super-Resolution With a Mosaic RGB Image
abstract
Recently, many hyperspectral (HS) image superresolution methods that merge a low spatial resolution HS image and a high spatial resolution three-channel RGB image have been proposed in spectral imaging. A largely ignored fact is that most existing commercial RGB cameras capture high resolution images by a single CCD/CMOS sensor equipped with a color filter array (CFA). In this paper, we account for the common imaging mechanism of commercial RGB cameras, and propose to use a mosaic RGB image for HS image super-resolution, which prevents demosaicing error and thus its propagation into the HS image super-resolution results. We design a proper nonlocal low-rank regularization to exploit the intrinsic properties - rich self-repeating patterns and high correlation across spectra - within HS images of natural scenes, and formulate the HS image super-resolution task into a variational optimization problem, which can be efficiently solved via the alternating direction method of multipliers (ADMM). The effectiveness of the proposed method has been evaluated on two benchmark datasets, demonstrating that the proposed method can provide substantial improvement over the current state-of-the-art HS image superresolution methods without considering the mosaicing effect. Finally, we show that our method can also perform well in the real capture system.
Ying Fu 0001, Yinqiang Zheng, Hua Huang 0001, Imari Sato, Yoichi Sato 0001
IEEE Trans. Image Process.5
2017 Rapid Prototyping of Accessible Interfaces With Gaze-Contingent Tunnel Vision Simulation
abstract
Active involvement of users with disabilities is difficult to employ during the iterative stages of the design process due to high costs and effort associated with user studies. This research proposes a user centered design (UCD) strategy to incorporate the use of gaze-contingent tunnel vision simulation with sighted individuals to facilitate rapid prototyping of accessible interfaces. Through three types of validation studies, we examined how our simulation techniques can provide the opportunity for continued evaluation and refinement of the design. Our simulation approach was effective in emulating scanning behaviors caused by tunnel vision along with grasping user feedback to recognize user interface and usability criteria early in the design cycle.
Rie Kamikubo, Keita Higuchi, Ryo Yonetani, Hideki Koike, Yoichi Sato 0001
ASSETS5
2017 EgoScanning: Quickly Scanning First-Person Videos with Egocentric Elastic Timelines
abstract
This work presents EgoScanning, a novel video fast-forwarding interface that helps users to find important events from lengthy first-person videos recorded with wearable cameras continuously. This interface is featured by an elastic timeline that adaptively changes playback speeds and emphasizes egocentric cues specific to first-person videos, such as hand manipulations, moving, and conversations with people, based on computer-vision techniques. The interface also allows users to input which of such cues are relevant to events of their interests. Through our user study, we confirm that users can find events of interests quickly from first-person videos thanks to the following benefits of using the EgoScanning interface: 1) adaptive changes of playback speeds allow users to watch fast-forwarded videos more easily; 2) Emphasized parts of videos can act as candidates of events actually significant to users; 3) Users are able to select relevant egocentric cues depending on events of their interests.
Keita Higuchi, Ryo Yonetani, Yoichi Sato 0001
CHI3
2017 Fast Multi-frame Stereo Scene Flow with Motion Segmentation
abstract
We propose a new multi-frame method for efficiently computing scene flow (dense depth and optical flow) and camera ego-motion for a dynamic scene observed from a moving stereo camera rig. Our technique also segments out moving objects from the rigid scene. In our method, we first estimate the disparity map and the 6-DOF camera motion using stereo matching and visual odometry. We then identify regions inconsistent with the estimated camera motion and compute per-pixel optical flow only at these regions. This flow proposal is fused with the camera motion-based flow proposal using fusion moves to obtain the final optical flow and motion segmentation. This unified framework benefits all four tasks - stereo, optical flow, visual odometry and motion segmentation leading to overall higher accuracy and efficiency. Our method is currently ranked third on the KITTI 2015 scene flow benchmark. Furthermore, our CPU implementation runs in 2-3 seconds per frame which is 1-3 orders of magnitude faster than the top six methods. We also report a thorough evaluation on challenging Sintel sequences with fast camera and object motion, where our method consistently outperforms OSF [30], which is currently ranked second on the KITTI benchmark.
Tatsunori Taniai, Sudipta N. Sinha, Yoichi Sato 0001
CVPR3
2017 From RGB to Spectrum for Natural Scenes via Manifold-Based Mapping
abstract
Spectral analysis of natural scenes can provide much more detailed information about the scene than an ordinary RGB camera. The richer information provided by hyperspectral images has been beneficial to numerous applications, such as understanding natural environmental changes and classifying plants and soils in agriculture based on their spectral properties. In this paper, we present an efficient manifold learning based method for accurately reconstructing a hyperspectral image from a single RGB image captured by a commercial camera with known spectral response. By applying a nonlinear dimensionality reduction technique to a large set of natural spectra, we show that the spectra of natural scenes lie on an intrinsically low dimensional manifold. This allows us to map an RGB vector to its corresponding hyperspectral vector accurately via our proposed novel manifold-based reconstruction pipeline. Experiments using both synthesized RGB images using hyperspectral datasets and real world data demonstrate our method outperforms the state-of-the-art.
Yan Jia 0005, Yinqiang Zheng, Lin Gu 0003, Art Subpa-Asa, Antony Lam, Yoichi Sato 0001, Imari Sato
ICCV6
2017 Privacy-Preserving Visual Learning Using Doubly Permuted Homomorphic Encryption
abstract
We propose a privacy-preserving framework for learning visual classifiers by leveraging distributed private image data. This framework is designed to aggregate multiple classifiers updated locally using private data and to ensure that no private information about the data is exposed during and after its learning procedure. We utilize a homomorphic cryptosystem that can aggregate the local classifiers while they are encrypted and thus kept secret. To overcome the high computational cost of homomorphic encryption of high-dimensional classifiers, we (1) impose sparsity constraints on local classifier updates and (2) propose a novel efficient encryption scheme named doublypermuted homomorphic encryption (DPHE) which is tailored to sparse high-dimensional data. DPHE (i) decomposes sparse data into its constituent non-zero values and their corresponding support indices, (ii) applies homomorphic encryption only to the non-zero values, and (iii) employs double permutations on the support indices to make them secret. Our experimental evaluation on several public datasets shows that the proposed approach achieves comparable performance against state-of-the-art visual recognition methods while preserving privacy and significantly outperforms other privacy-preserving methods.
Ryo Yonetani, Vishnu Naresh Boddeti, Kris Makoto Kitani, Yoichi Sato 0001
ICCV4
2017 Community detection using random-walk similarity and application to image clustering
abstract
The technology used to detect community structures in graphs, or graph clustering technology, is important in a wide range of disciplines, such as sociology, biology, and computer science. Previously, many successful community detection methods have relied on the optimization of a quantity referred to as modularity, which is a quality index for the partition of a graph into communities. However, such methods suffer from a key drawback, namely, the inability to identify relatively small communities. To overcome this drawback, we propose a novel community detection method that can detect small communities. This is based on the property that a random walker will not readily leave a community even if it is small. The work presented in this paper demonstrates that our method detects both small and large communities in the practical application of clustering tourist attraction images obtained from Flickr.
Makoto Okuda, Shin'ichi Satoh 0001, Shoichiro Iwasawa, Shunsuke Yoshida, Yutaka Kidawara, Yoichi Sato 0001
ICIP6
2017 Adaptive Spatial-Spectral Dictionary Learning for Hyperspectral Image Restoration
Ying Fu 0001, Antony Lam, Imari Sato, Yoichi Sato 0001
Int. J. Comput. Vis.4
2017 An Ego-Vision System for Hand Grasp Analysis
abstract
This paper presents an egocentric vision (ego-vision) system for hand grasp analysis in unstructured environments. Our goal is to automatically recognize hand grasp types and to discover the visual structures of hand grasps using a wearable camera. In the proposed system, free hand–object interactions are recorded from a first-person viewing perspective. State-of-the-art computer vision techniques are used to detect hands and extract hand-based features. A new feature representation that incorporates hand tracking information is also proposed. Then, grasp classifiers are trained to discriminate among different grasp types from a predefined grasp taxonomy. Based on the trained grasp classifiers, visual structures of hand grasps are learned using an iterative grasp clustering method. In experiments, grasp recognition performance in both laboratory and real-world scenarios is evaluated. The best classification accuracy our system achieves is $\text{92}\%$ and $\text{59}\%$ , respectively. System generality to different tasks and users is also verified by the experiments. Analysis in a real-world scenario shows that it is possible to automatically learn intuitive visual grasp structures that are consistent with expert-designed grasp taxonomies.
Minjie Cai, Kris Makoto Kitani, Yoichi Sato 0001
IEEE Trans. Hum. Mach. Syst.3
2017 Appearance-Based Gaze Estimation via Uncalibrated Gaze Pattern Recovery
abstract
Aiming at reducing the restrictions due to person/scene dependence, we deliver a novel method that solves appearance-based gaze estimation in a novel fashion. First, we introduce and solve an "uncalibrated gaze pattern" solely from eye images independent of the person and scene. The gaze pattern recovers gaze movements up to only scaling and translation ambiguities, via nonlinear dimension reduction and pixel motion analysis, while no training/calibration is needed. This is new in the literature and enables novel applications. Second, our method allows simple calibrations to align the gaze pattern to any gaze target. This is much simpler than conventional calibrations which rely on sufficient training data to compute person and scene-specific nonlinear gaze mappings. Through various evaluations, we show that: 1) the proposed uncalibrated gaze pattern has novel and broad capabilities; 2) the proposed calibration is simple and efficient, and can be even omitted in some scenarios; and 3) quantitative evaluations produce promising results under various conditions.
Feng Lu 0005, Xiaowu Chen 0001, Yoichi Sato 0001
IEEE Trans. Image Process.3
2016 Visual Guidance with Unnoticed Blur Effect
abstract
In information media such as TV programs, digital signage, or web pages, information content providers often want to guide viewers' attention to a particular location of the display. However, "active" methods, such as flashing displays, using animation, or changing colors, often interrupt viewers' concentration and makes viewers feel annoyed. This paper proposes a method for guiding viewers' attention without viewers noticing. By focusing on a characteristic of the human visual system, we propose a dynamic blur control method. Our method gradually blurs the image on the display to the threshold at which viewers are aware of the modulation of the display, while the region where viewers' attention should be guided remains unblurred. Two subjective experiments were conducted to show the effectiveness of our method. In the first, viewers' attention was guided to the unblurred region using blur control. In the second, a threshold was found at which viewers were aware of the modulation, and viewers' gaze is guided below this threshold. This means that the viewers' attention can be guided without them noticing.
Hajime Hata, Hideki Koike, Yoichi Sato 0001
AVI3
2016 Can Eye Help You?: Effects of Visualizing Eye Fixations on Remote Collaboration Scenarios for Physical Tasks
abstract
In this work, we investigate how remote collaboration between a local worker and a remote collaborator will change if eye fixations of the collaborator are presented to the worker. We track the collaborator's points of gaze on a monitor screen displaying a physical workspace and visualize them onto the space by a projector or through an optical see-through head-mounted display. Through a series of user studies, we have found the followings: 1) Eye fixations can serve as a fast and precise pointer to objects of the collaborator's interest. 2) Eyes and other modalities, such as hand gestures and speech, are used differently for object identification and manipulation. 3) Eyes are used for explicit instructions only when they are combined with speech. 4) The worker can predict some intentions of the collaborator such as his/her current interest and next instruction.
Keita Higuchi, Ryo Yonetani, Yoichi Sato 0001
CHI3
2016 Exploiting Spectral-Spatial Correlation for Coded Hyperspectral Image Restoration
abstract
Conventional scanning and multiplexing techniques for hyperspectral imaging suffer from limited temporal and/or spatial resolution. To resolve this issue, coding techniques are becoming increasingly popular in developing snapshot systems for high-resolution hyperspectral imaging. For such systems, it is a critical task to accurately restore the 3D hyperspectral image from its corresponding coded 2D image. In this paper, we propose an effective method for coded hyperspectral image restoration, which exploits extensive structure sparsity in the hyperspectral image. Specifically, we simultaneously explore spectral and spatial correlation via low-rank regularizations, and formulate the restoration problem into a variational optimization model, which can be solved via an iterative numerical algorithm. Experimental results using both synthetic data and real images show that the proposed method can significantly outperform the state-of-the-art methods on several popular coding-based hyperspectral imaging systems.
Ying Fu 0001, Yinqiang Zheng, Imari Sato, Yoichi Sato 0001
CVPR4
2016 Hierarchical Gaussian Descriptor for Person Re-identification
abstract
Describing the color and textural information of a person image is one of the most crucial aspects of person re-identification. In this paper, we present a novel descriptor based on a hierarchical distribution of pixel features. A hierarchical covariance descriptor has been successfully applied for image classification. However, the mean information of pixel features, which is absent in covariance, tends to be major discriminative information of person images. To solve this problem, we describe a local region in an image via hierarchical Gaussian distribution in which both means and covariances are included in their parameters. More specifically, we model the region as a set of multiple Gaussian distributions in which each Gaussian represents the appearance of a local patch. The characteristics of the set of Gaussians are again described by another Gaussian distribution. In both steps, unlike the hierarchical covariance descriptor, the proposed descriptor can model both the mean and the covariance information of pixel features properly. The results of experiments conducted on five databases indicate that the proposed descriptor exhibits remarkably high performance which outperforms the state-of-the-art descriptors for person re-identification.
Tetsu Matsukawa, Takahiro Okabe, Einoshin Suzuki, Yoichi Sato 0001
CVPR4
2016 Joint Recovery of Dense Correspondence and Cosegmentation in Two Images
abstract
We propose a new technique to jointly recover cosegmentation and dense per-pixel correspondence in two images. Our method parameterizes the correspondence field using piecewise similarity transformations and recovers a mapping between the estimated common "foreground" regions in the two images allowing them to be precisely aligned. Our formulation is based on a hierarchical Markov random field model with segmentation and transformation labels. The hierarchical structure uses nested image regions to constrain inference across multiple scales. Unlike prior hierarchical methods which assume that the structure is given, our proposed iterative technique dynamically recovers the structure along with the labeling. This joint inference is performed in an energy minimization framework using iterated graph cuts. We evaluate our method on a new dataset of 400 image pairs with manually obtained ground truth, where it outperforms state-of-the-art methods designed specifically for either cosegmentation or correspondence estimation.
Tatsunori Taniai, Sudipta N. Sinha, Yoichi Sato 0001
CVPR3
2016 Recognizing Micro-Actions and Reactions from Paired Egocentric Videos
abstract
We aim to understand the dynamics of social interactions between two people by recognizing their actions and reactions using a head-mounted camera. Our work will impact several first-person vision tasks that need the detailed understanding of social interactions, such as automatic video summarization of group events and assistive systems. To recognize micro-level actions and reactions, such as slight shifts in attention, subtle nodding, or small hand actions, where only subtle body motion is apparent, we propose to use paired egocentric videos recorded by two interacting people. We show that the first-person and second-person points-of-view features of two people, enabled by paired egocentric videos, are complementary and essential for reliably recognizing micro-actions and reactions. We also build a new dataset of dyadic (two-persons) interactions that comprises more than 1000 pairs of egocentric videos to enable systematic evaluations on the task of micro-action and reaction recognition.
Ryo Yonetani, Kris Makoto Kitani, Yoichi Sato 0001
CVPR3
2016 Visual Motif Discovery via First-Person Vision
Ryo Yonetani, Kris Makoto Kitani, Yoichi Sato 0001
ECCV (2)3
2016 Special issue on IAPR MVA2013 best papers
Masaki Suwa, Yoichi Sato 0001, Katsushi Ikeuchi
Mach. Vis. Appl.2
2016 Separating Reflective and Fluorescent Components Using High Frequency Illumination in the Spectral Domain
abstract
Hyperspectral imaging is beneficial to many applications but most traditional methods do not consider fluorescent effects which are present in everyday items ranging from paper to even our food. Furthermore, everyday fluorescent items exhibit a mix of reflection and fluorescence so proper separation of these components is necessary for analyzing them. In recent years, effective imaging methods have been proposed but most require capturing the scene under multiple illuminants. In this paper, we demonstrate efficient separation and recovery of reflectance and fluorescence emission spectra through the use of two high frequency illuminations in the spectral domain. With the obtained fluorescence emission spectra from our high frequency illuminants, we then describe how to estimate the fluorescence absorption spectrum of a material given its emission spectrum. In addition, we provide an in depth analysis of our method and also show that filters can be used in conjunction with standard light sources to generate the required high frequency illuminants. We also test our method under ambient light and demonstrate an application of our method to synthetic relighting of real scenes.
Ying Fu 0001, Antony Lam, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2016 Reflectance and Fluorescence Spectral Recovery via Actively Lit RGB Images
abstract
In recent years, fluorescence analysis of scenes has received attention in computer vision. Fluorescence can provide additional information about scenes, and has been used in applications such as camera spectral sensitivity estimation, 3D reconstruction, and color relighting. In particular, hyperspectral images of reflective-fluorescent scenes provide a rich amount of data. However, due to the complex nature of fluorescence, hyperspectral imaging methods rely on specialized equipment such as hyperspectral cameras and specialized illuminants. In this paper, we propose a more practical approach to hyperspectral imaging of reflective-fluorescent scenes using only a conventional RGB camera and varied colored illuminants. The key idea of our approach is to exploit a unique property of fluorescence: the chromaticity of fluorescent emissions are invariant under different illuminants. This allows us to robustly estimate spectral reflectance and fluorescent emission chromaticity. We then show that given the spectral reflectance and fluorescent chromaticity, the fluorescence absorption and emission spectra can also be estimated. We demonstrate in results that all scene spectra can be accurately estimated from RGB images. Finally, we show that our method can be used to accurately relight scenes under novel lighting.
Ying Fu 0001, Antony Lam, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2015 Uncalibrated photometric stereo based on elevation angle recovery from BRDF symmetry of isotropic materials
abstract
This paper addresses the problem of uncalibrated photometric stereo with isotropic reflectances. Existing methods face difficulty in solving for the elevation angles of surface normals when the light sources only cover the visible hemisphere. Here, we introduce the notion of “constrained half-vector symmetry” for general isotropic BRDFs and show its capability of elevation angle recovery. This sort of symmetry can be observed in a 1D BRDF slice from a subset of surface normals with the same azimuth angle, and we use it to devise an efficient modeling and solution method to constrain and recover the elevation angles of surface normals accurately. To enable our method to work in an uncalibrated manner, we further solve for light sources in the case of general isotropic BRDFs. By combining this method with the existing ones for azimuth angle estimation, we can get state-of-the-art results for uncalibrated photometric stereo with general isotropic reflectances.
Feng Lu 0005, Imari Sato, Yoichi Sato 0001
CVPR3
2015 Ego-surfing first person videos
abstract
We envision a future time when wearable cameras (e.g., small cameras in glasses or pinned on a shirt collar) are worn by the masses and record first-person point-of-view (POV) videos of everyday life. While these cameras can enable new assistive technologies and novel research challenges, they also raise serious privacy concerns. For example, first-person videos passively recorded by wearable cameras will necessarily include anyone who comes into the view of a camera - with or without consent. Motivated by these benefits and risks, we develop a self-search technique tailored to first-person POV videos. The key observation of our work is that the egocentric head motions of a target person (i.e., the self) are observed both in the POV video of the target and observer. The motion correlation between the target person's video and the observer's video can then be used to uniquely identify instances of the self. We incorporate this feature into our proposed approach that computes the motion correlation over supervoxel hierarchies to localize target instances in observer videos. Our proposed approach significantly improves self-search performance over several well-known face detectors and recognizers. Furthermore, we show how our approach can enable several practical applications such as privacy filtering, automated video collection and social group discovery.
Ryo Yonetani, Kris Makoto Kitani, Yoichi Sato 0001
CVPR3
2015 Illumination and reflectance spectra separation of a hyperspectral image meets low-rank matrix factorization
abstract
This paper addresses the illumination and reflectance spectra separation (IRSS) problem of a hyperspectral image captured under general spectral illumination. The huge amount of pixels in a hypersepctral image poses tremendous challenges on computational efficiency, yet in turn offers greater color variety that might be utilized to improve separation accuracy and relax the restrictive subspace illumination assumption in existing works. We show that this IRSS problem can be modeled into a low-rank matrix factorization problem, and prove that the separation is unique up to an unknown scale under the standard low-dimensionality assumption of reflectance. We also develop a scalable algorithm for this separation task that works in the presence of model error and image noise. Experiments on both synthetic data and real images have demonstrated that our separation results are sufficiently accurate, and can benefit some important applications, such as spectra relighting and illumination swapping.
Yinqiang Zheng, Imari Sato, Yoichi Sato 0001
CVPR3
2015 Adaptive Spatial-Spectral Dictionary Learning for Hyperspectral Image Denoising
abstract
Hyperspectral imaging is beneficial in a diverse range of applications from diagnostic medicine, to agriculture, to surveillance to name a few. However, hyperspectral images often times suffer from degradation due to the limited light, which introduces noise into the imaging process. In this paper, we propose an effective model for hyperspectral image (HSI) denoising that considers underlying characteristics of HSIs: sparsity across the spatial-spectral domain, high correlation across spectra, and non-local self-similarity over space. We first exploit high correlation across spectra and non-local self-similarity over space in the noisy HSI to learn an adaptive spatial-spectral dictionary. Then, we employ the local and non-local sparsity of the HSI under the learned spatial-spectral dictionary to design an HSI denoising model, which can be effectively solved by an iterative numerical algorithm with parameters that are adaptively adjusted for different clusters and different noise levels. Experimental results on HSI denoising show that the proposed method can provide substantial improvements over the current state-of-the-art HSI denoising methods in terms of both objective metric and subjective visual quality.
Ying Fu 0001, Antony Lam, Imari Sato, Yoichi Sato 0001
ICCV4
2015 Separating Fluorescent and Reflective Components by Using a Single Hyperspectral Image
abstract
This paper introduces a novel method to separate fluorescent and reflective components in the spectral domain. In contrast to existing methods, which require to capture two or more images under varying illuminations, we aim to achieve this separation task by using a single hyperspectral image. After identifying the critical hurdle in single-image component separation, we mathematically design the optimal illumination spectrum, which is shown to contain substantial high-frequency components in the frequency domain. This observation, in turn, leads us to recognize a key difference between reflectance and fluorescence in response to the frequency modulation effect of illumination, which fundamentally explains the feasibility of our method. On the practical side, we successfully find an off-the-shelf lamp as the light source, which is strong in irradiance intensity and cheap in cost. A fast linear separation algorithm is developed as well. Experiments using both synthetic data and real images have confirmed the validity of the selected illuminant and the accuracy of our separation algorithm.
Yinqiang Zheng, Ying Fu 0001, Antony Lam, Imari Sato, Yoichi Sato 0001
ICCV5
2015 Fast sparse edge-based intrinsic image decomposition guided by chromaticity gradients
abstract
In this paper, we proposed a novel edge-based method for intrinsic image decomposition. To be specific, to address this ill-posed problem via single image, we employed the information of chromaticity gradient to guide recovering the reflectance gradient map, and regularized the Retinex assumption by the ℓ0-norm to render the reflectance layer piece-wise smooth. Our method is simple but effective. Qualitative together with quantitative experiments were carried out on some public datasets. The comparison results show that our method runs much faster than the state-of-art methods with comparable performance.
Jinze Yu 0003, Yoichi Sato 0001
ICIP2
2015 A scalable approach for understanding the visual structures of hand grasps
abstract
Our goal is to automatically recognize hand grasps and to discover the visual structures (relationships) between hand grasps using wearable cameras. Wearable cameras provide a first-person perspective which enables continuous visual hand grasp analysis of everyday activities. In contrast to previous work focused on manual analysis of first-person videos of hand grasps, we propose a fully automatic vision-based approach for grasp analysis. A set of grasp classifiers are trained for discriminating between different grasp types based on large margin visual predictors. Building on the output of these grasp classifiers, visual structures among hand grasps are learned based on an iterative discriminative clustering procedure. We first evaluated our classifiers on a controlled indoor grasp dataset and then validated the analytic power of our approach on real-world data taken from a machinist. The average F1 score of our grasp classifiers achieves over 0.80 for the indoor grasp dataset. Analysis of real-world video shows that it is possible to automatically learn intuitive visual grasp structures that are consistent with expert-designed grasp taxonomies.
Minjie Cai, Kris Makoto Kitani, Yoichi Sato 0001
ICRA3
2015 Editorial
Björn Stenger, Norimichi Ukita, Yoichi Sato 0001, Pascal Fua, David J. Fleet
Comput. Vis. Image Underst.3
2015 From Intensity Profile to Surface Normal: Photometric Stereo for Unknown Light Sources and Isotropic Reflectances
abstract
We propose an uncalibrated photometric stereo method that works with general and unknown isotropic reflectances. Our method uses a pixel intensity profile, which is a sequence of radiance intensities recorded at a pixel under unknown varying directional illumination. We show that for general isotropic materials and uniformly distributed light directions, the geodesic distance between intensity profiles is linearly related to the angular difference of their corresponding surface normals, and that the intensity distribution of the intensity profile reveals reflectance properties. Based on these observations, we develop two methods for surface normal estimation; one for a general setting that uses only the recorded intensity profiles, the other for the case where a BRDF database is available while the exact BRDF of the target scene is still unknown. Quantitative and qualitative evaluations are conducted using both synthetic and real-world scenes, which show the state-of-the-art accuracy of smaller than 10 degree without using reference data and 5 degree with reference data for all 100 materials in MERL database.
Feng Lu 0005, Yasuyuki Matsushita, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2015 Appearance-Based Gaze Estimation With Online Calibration From Mouse Operations
abstract
This paper presents an unconstrained gaze estimation method using an online learning algorithm. We focus on a desktop scenario, where a user operates a personal computer, and use the mouse-clicked positions to infer, where on the screen the user is looking at. Our method continuously captures the user's head pose and eye images with a monocular camera, and each mouse click triggers learning sample acquisition. In order to handle head pose variations, the samples are adaptively clustered according to the estimated head pose. Then, local reconstruction-based gaze estimation models are incrementally updated in each cluster. We conducted a prototype evaluation in real-world environments, and our method achieved an estimation accuracy of 2.9°.
Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001, Hideki Koike
IEEE Trans. Hum. Mach. Syst.3
2015 Gaze Estimation From Eye Appearance: A Head Pose-Free Method via Eye Image Synthesis
abstract
In this paper, we address the problem of free head motion in appearance-based gaze estimation. This problem remains challenging because head motion changes eye appearance significantly, and thus, training images captured for an original head pose cannot handle test images captured for other head poses. To overcome this difficulty, we propose a novel gaze estimation method that handles free head motion via eye image synthesis based on a single camera. Compared with conventional fixed head pose methods with original training images, our method only captures four additional eye images under four reference head poses, and then, precisely synthesizes new training images for other unseen head poses in estimation. To this end, we propose a single-directional (SD) flow model to efficiently handle eye image variations due to head motion. We show how to estimate SD flows for reference head poses first, and then use them to produce new SD flows for training image synthesis. Finally, with synthetic training images, joint optimization is applied that simultaneously solves an eye image alignment and a gaze estimation. Evaluation of the method was conducted through experiments to assess its performance and demonstrate its effectiveness.
Feng Lu 0005, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001
IEEE Trans. Image Process.4
2015 Cell Detection From Redundant Candidate Regions Under Nonoverlapping Constraints
abstract
Cell detection in microscopy images is essential for automated cell behavior analysis including cell shape analysis and cell tracking. Robust cell detection in high-density and low-contrast images is still challenging since cells often touch and partially overlap, forming a cell cluster with blurry intercellular boundaries. In such cases, current methods tend to detect multiple cells as a cluster. If the control parameters are adjusted to separate the touching cells, other problems often occur: a single cell may be segmented into several regions, and cells in low-intensity regions may not be detected. To solve these problems, we first detect redundant candidate regions, which include many false positives but in turn very few false negatives, by allowing candidate regions to overlap with each other. Next, the score for how likely the candidate region contains the main part of a single cell is computed for each cell candidate using supervised learning. Then we select an optimal set of cell regions from the redundant regions under nonoverlapping constraints, where each selected region looks like a single cell and the selected regions do not overlap. We formulate this problem of optimal region selection as a binary linear programming problem under nonoverlapping constraints. We demonstrated the effectiveness of our method for several types of cells in microscopy images. Our method performed better than five representative methods, achieving an F-measure of over 0.9 for all data sets. Experimental application of the proposed method to 3-D images demonstrated that also works well for 3-D cell detection.
Ryoma Bise, Yoichi Sato 0001
IEEE Trans. Medical Imaging2
2014 Shape-Preserving Half-Projective Warps for Image Stitching
abstract
This paper proposes a novel parametric warp which is a spatial combination of a projective transformation and a similarity transformation. Given the projective transformation relating two input images, based on an analysis of the projective transformation, our method smoothly extrapolates the projective transformation of the overlapping regions into the non-overlapping regions and the resultant warp gradually changes from projective to similarity across the image. The proposed warp has the strengths of both projective and similarity warps. It provides good alignment accuracy as projective warps while preserving the perspective of individual image as similarity warps. It can also be combined with more advanced local-warp-based alignment methods such as the as-projective-as-possible warp for better alignment accuracy. With the proposed warp, the field of view can be extended by stitching images with less projective distortion (stretched shapes and enlarged sizes).
Che-Han Chang, Yoichi Sato 0001, Yung-Yu Chuang
CVPR2
2014 Reflectance and Fluorescent Spectra Recovery Based on Fluorescent Chromaticity Invariance under Varying Illumination
abstract
In recent years, fluorescence analysis of scenes has received attention. Fluorescence can provide additional information about scenes, and has been used in applications such as camera spectral sensitivity estimation, 3D reconstruction, and color relighting. In particular, hyperspectral images of reflective-fluorescent scenes provide a rich amount of data. However, due to the complex nature of fluorescence, hyperspectral imaging methods rely on specialized equipment such as hyperspectral cameras and specialized illuminants. In this paper, we propose a more practical approach to hyperspectral imaging of reflective-fluorescent scenes using only a conventional RGB camera and varied colored illuminants. The key idea of our approach is to exploit a unique property of fluorescence: the chromaticity of fluorescence emissions are invariant under different illuminants. This allows us to robustly estimate spectral reflectance and fluorescence emission chromaticity. We then show that given the spectral reflectance and fluorescent chromaticity, the fluorescence absorption and emission spectra can also be estimated. We demonstrate in results that all scene spectra can be accurately estimated from RGB images. Finally, we show that our method can be used to accurately relight scenes under novel lighting.
Ying Fu 0001, Antony Lam, Yasuyuki Kobashi, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
CVPR6
2014 Learning-by-Synthesis for Appearance-Based 3D Gaze Estimation
abstract
Inferring human gaze from low-resolution eye images is still a challenging task despite its practical importance in many application scenarios. This paper presents a learning-by-synthesis approach to accurate image-based gaze estimation that is person- and head pose-independent. Unlike existing appearance-based methods that assume person-specific training data, we use a large amount of cross-subject training data to train a 3D gaze estimator. We collect the largest and fully calibrated multi-view gaze dataset and perform a 3D reconstruction in order to generate dense training data of eye images. By using the synthesized dataset to learn a random regression forest, we show that our method outperforms existing methods that use low-resolution eye images.
Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001
CVPR3
2014 Interreflection Removal Using Fluorescence
Ying Fu 0001, Antony Lam, Yasuyuki Matsushita, Imari Sato, Yoichi Sato 0001
ECCV (5)5
2014 Spectra Estimation of Fluorescent and Reflective Scenes by Using Ordinary Illuminants
Yinqiang Zheng, Imari Sato, Yoichi Sato 0001
ECCV (5)3
2014 Influence of stimulus and viewing task types on a learning-based visual saliency model
abstract
Learning-based approaches using actual human gaze data have been proven to be an efficient way to acquire accurate visual saliency models and attracted much interest in recent years. However, it still remains yet to be answered how different types of stimulus (e.g., fractal images, and natural images with or without human faces) and viewing tasks (e.g., free viewing or a preference rating task) affect learned visual saliency models. In this study, we quantitatively investigate how learned saliency models differ when using datasets collected in different settings (image contextual level and viewing task) and discuss the importance of choosing appropriate experimental settings.
Binbin Ye, Yusuke Sugano, Yoichi Sato 0001
ETRA3
2014 Person Re-identification via Discriminative Accumulation of Local Features
abstract
Metric learning to learn a good distance metric for distinguishing different people while being insensitive to intra-person variations is widely applied to person re-identification. In previous works, local histograms are densely sampled to extract spatially localized information of each person image. The extracted local histograms are then concatenated into one vector that is used as an input of metric learning. However, the dimensionality of such a concatenated vector often becomes large while the number of training samples is limited. This leads to an over fitting problem. In this work, we argue that such a problem of over-fitting comes from that it is each local histogram dimension (e.g. color brightness bin) in the same position is treated separately to examine which part of the image is more discriminative. To solve this problem, we propose a method that analyzes discriminative image positions shared by different local histogram dimensions. A common weight map shared by different dimensions and a distance metric which emphasizes discriminative dimensions in the local histogram are jointly learned with a unified discriminative criterion. Our experiments using four different public datasets confirmed the effectiveness of the proposed method.
Tetsu Matsukawa, Takahiro Okabe, Yoichi Sato 0001
ICPR3
2014 Fast Spectral Reflectance Recovery Using DLP Projector
Shuai Han 0006, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
Int. J. Comput. Vis.4
2014 Learning gaze biases with head motion for head pose-free gaze estimation
Feng Lu 0005, Takahiro Okabe, Yusuke Sugano, Yoichi Sato 0001
Image Vis. Comput.4
2014 Adaptive Linear Regressionfor Appearance-Based Gaze Estimation
abstract
We investigate the appearance-based gaze estimation problem, with respect to its essential difficulty in reducing the number of required training samples, and other practical issues such as slight head motion, image resolution variation, and eye blinking. We cast the problem as mapping high-dimensional eye image features to low-dimensional gaze positions, and propose an adaptive linear regression (ALR) method as the key to our solution. The ALR method adaptively selects an optimal set of sparsest training samples for the gaze estimation via ℓ(1)-optimization. In this sense, the number of required training samples is significantly reduced for high accuracy estimation. In addition, by adopting the basic ALR objective function, we integrate the gaze estimation, subpixel alignment and blink detection into a unified optimization framework. By solving these problems simultaneously, we successfully handle slight head motion, image resolution variation and eye blinking in appearance-based gaze estimation. We evaluated the proposed method by conducting experiments with multiple users and variant conditions to verify its effectiveness.
Feng Lu 0005, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2013 Social Group Discovery from Surveillance Videos: A Data-Driven Approach with Attention-Based Cues
abstract
This paper presents an approach to discover social groups in surveillance videos by incorporating attention-based cues to model group behaviors of pedestrians in videos. Group behaviors are modeled as a set of decision trees with the decisions being basic measurements based on positionbased and attention-based cues. Rather than enforcing explicit models, we apply tree-based learning algorithms to implicitly construct the decision tree models. The experimental results demonstrate that incorporating attention-based cues significantly increased the estimation accuracy compared to the conventional approaches that used position-based cues alone.
Isarun Chamveha, Yusuke Sugano, Yoichi Sato 0001, Akihiro Sugimoto
BMVC3
2013 Spectral Imaging Using Basis Lights
abstract
Antony Lam1 http://research.nii.ac.jp/~antony Art Subpa-Asa2 [email protected] Imari Sato1 http://research.nii.ac.jp/~imarik Takahiro Okabe3 http://www.pluto.ai.kyutech.ac.jp/~okabe Yoichi Sato4 http://www.hci.iis.u-tokyo.ac.jp/~ysato 1 National Institute of Informatics Tokyo, Japan 2 The Stock Exchange of Thailand Bangkok, Thailand 3 Kyushu Institute of Technology Fukuoka, Japan 4 The University of Tokyo Tokyo, Japan
Antony Lam, Art Subpa-Asa, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
BMVC5
2013 Uncalibrated Photometric Stereo for Unknown Isotropic Reflectances
abstract
We propose an uncalibrated photometric stereo method that works with general and unknown isotropic reflectances. Our method uses a pixel intensity profile, which is a sequence of radiance intensities recorded at a pixel across multi-illuminance images. We show that for general isotropic materials, the geodesic distance between intensity profiles is linearly related to the angular difference of their surface normals, and that the intensity distribution of an intensity profile conveys information about the reflectance properties, when the intensity profile is obtained under uniformly distributed directional lightings. Based on these observations, we show that surface normals can be estimated up to a convex/concave ambiguity. A solution method based on matrix decomposition with missing data is developed for a reliable estimation. Quantitative and qualitative evaluations of our method are performed using both synthetic and real-world scenes.
Feng Lu 0005, Yasuyuki Matsushita, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
CVPR5
2013 Separating Reflective and Fluorescent Components Using High Frequency Illumination in the Spectral Domain
abstract
Hyper spectral imaging is beneficial to many applications but current methods do not consider fluorescent effects which are present in everyday items ranging from paper, to clothing, to even our food. Furthermore, everyday fluorescent items exhibit a mix of reflectance and fluorescence. So proper separation of these components is necessary for analyzing them. In this paper, we demonstrate efficient separation and recovery of reflective and fluorescent emission spectra through the use of high frequency illumination in the spectral domain. With the obtained fluorescent emission spectra from our high frequency illuminants, we then present to our knowledge, the first method for estimating the fluorescent absorption spectrum of a material given its emission spectrum. Conventional bispectral measurement of absorption and emission spectra needs to examine all combinations of incident and observed light wavelengths. In contrast, our method requires only two hyper spectral images. The effectiveness of our proposed methods are then evaluated through a combination of simulation and real experiments. We also demonstrate an application of our method to synthetic relighting of real scenes.
Ying Fu 0001, Antony Lam, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
ICCV5
2013 Head direction estimation from low resolution images with scene adaptation
Isarun Chamveha, Yusuke Sugano, Daisuke Sugimura, Teera Siriteerakul, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto
Comput. Vis. Image Underst.6
2013 Appearance-Based Gaze Estimation Using Visual Saliency
abstract
We propose a gaze sensing method using visual saliency maps that does not need explicit personal calibration. Our goal is to create a gaze estimator using only the eye images captured from a person watching a video clip. Our method treats the saliency maps of the video frames as the probability distributions of the gaze points. We aggregate the saliency maps based on the similarity in eye images to efficiently identify the gaze points from the saliency maps. We establish a mapping between the eye images to the gaze points by using Gaussian process regression. In addition, we use a feedback loop from the gaze estimator to refine the gaze probability maps to improve the accuracy of the gaze estimation. The experimental results show that the proposed method works well with different people and video clips and achieves a 3.5-degree accuracy, which is sufficient for estimating a user's attention on a display.
Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Graph-based joint clustering of fixations and visual entities
abstract
We present a method that extracts groups of fixations and image regions for the purpose of gaze analysis and image understanding. Since the attentional relationship between visual entities conveys rich information, automatically determining the relationship provides us a semantic representation of images. We show that, by jointly clustering human gaze and visual entities, it is possible to build meaningful and comprehensive metadata that offer an interpretation about how people see images. To achieve this, we developed a clustering method that uses a joint graph structure between fixation points and over-segmented image regions to ensure a cross-domain smoothness constraint. We show that the proposed clustering method achieves better performance in relating attention to visual entities in comparison with standard clustering techniques.
Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001
ACM Trans. Appl. Percept.3
2012 Toward Efficient Acquisition of BRDFs with Fewer Samples
Muhammad Asad Ali, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
ACCV (4)4
2012 Deblurring Vein Images and Removing Skin Wrinkle Patterns by Using Tri-band Illumination
Naoto Miura, Yoichi Sato 0001
ACCV (4)2
2012 Camera spectral sensitivity estimation from a single image under unknown illumination by using fluorescence
abstract
Camera spectral sensitivity plays an important role for various color-based computer vision tasks. Although several methods have been proposed to estimate it, their applicability is severely restricted by the requirement for a known illumination spectrum. In this work, we present a single-image estimation method using fluorescence with no requirement for a known illumination spectrum. Under different illuminations, the spectral distributions of fluorescence emitted from the same material remain unchanged up to a certain scale. Thus, a camera's response to the fluorescence would have the same chromaticity. Making use of this chromaticity invariance, the camera spectral sensitivity can be estimated under an arbitrary illumination whose spectrum is unknown. Through extensive experiments, we proved that our method is accurate under different illuminations. Moreover, we show how to recover the spectra of daylight from the estimated results. Finally, we use the estimated camera spectral sensitivities and daylight spectra to solve color correction problems.
Shuai Han 0006, Yasuyuki Matsushita, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
CVPR5
2012 Bispectral photometric stereo based on fluorescence
abstract
We propose a novel technique called bispectral photometric stereo that makes effective use of fluorescence for shape reconstruction. Fluorescence is a common phenomenon occurring in many objects from natural gems and corals, to fluorescent dyes used in clothing. One of the important characteristics of fluorescence is its wavelength-shifting behavior: fluorescent materials absorb light at a certain wavelength and then reemit it at longer wavelengths. Due to the complexity of its emission process, fluorescence tends to be excluded from most algorithms in computer vision and image processing. In this paper, we show that there is a strong similarity between fluorescence and ideal diffuse reflection and that fluorescence can provide distinct clues on how to estimate an object's shape. Moreover, fluorescence's wavelength-shifting property enables us to estimate the shape of an object by applying photometric stereo to emission-only images without suffering from specular reflection. This is the significant advantage of the fluorescence-based method over previous methods based on reflection.
Imari Sato, Takahiro Okabe, Yoichi Sato 0001
CVPR3
2012 Incorporating visual field characteristics into a saliency map
abstract
Characteristics of the human visual field are well known to be different in central (fovea) and peripheral areas. Existing computational models of visual saliency, however, do not take into account this biological evidence. The existing models compute visual saliency uniformly over the retina and, thus, have difficulty in accurately predicting the next gaze (fixation) point. This paper proposes to incorporate human visual field characteristics into visual saliency, and presents a computational model for producing such a saliency map. Our model integrates image features obtained by bottom-up computation in such a way that weights for the integration depend on the distance from the current gaze point where the weights are optimally learned using actual saccade data. The experimental results using a large number of fixation/saccade data with wide viewing angles demonstrate the advantage of our saliency map, showing that it can accurately predict the point where one looks next.
Hideyuki Kubota, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto, Kazuo Hiraki
ETRA4
2012 Denoising hyperspectral images using spectral domain statistics
Antony Lam, Imari Sato, Yoichi Sato 0001
ICPR3
2012 Head pose-free appearance-based gaze sensing via eye image synthesis
Feng Lu 0005, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001
ICPR4
2012 Illumination normalization of face images with cast shadows
Tetsu Matsukawa, Takahiro Okabe, Yoichi Sato 0001
ICPR3
2011 A Head Pose-free Approach for Appearance-based Gaze Estimation
abstract
To infer human gaze from eye appearance, various methods have been proposed. However, most of them assume a fixed head pose because allowing free head motion adds 6 degrees of freedom to the problem and requires a prohibitively large number of training samples. In this paper, we aim at solving the appearance-based gaze estimation problem under free head motion without significantly increasing the cost of training. The idea is to decompose the problem into subproblems, including initial estimation under fixed head pose and subsequent compensations for estimation biases caused by head rotation and eye appearance distortion. Then each subproblem is solved by either learning-based method or geometric-based calculation. Specifically, the gaze estimation bias caused by eye appearance distortion is learnt effectively from a 5-seconds video clip. Extensive experiments were conducted to verify the effectiveness of the proposed approach. 1
Feng Lu 0005, Takahiro Okabe, Yusuke Sugano, Yoichi Sato 0001
BMVC4
2011 Fast unsupervised ego-action learning for first-person sports videos
abstract
Portable high-quality sports cameras (e.g. head or helmet mounted) built for recording dynamic first-person video footage are becoming a common item among many sports enthusiasts. We address the novel task of discovering first-person action categories (which we call ego-actions) which can be useful for such tasks as video indexing and retrieval. In order to learn ego-action categories, we investigate the use of motion-based histograms and unsupervised learning algorithms to quickly cluster video content. Our approach assumes a completely unsupervised scenario, where labeled training videos are not available, videos are not pre-segmented and the number of ego-action categories are unknown. In our proposed framework we show that a stacked Dirichlet process mixture model can be used to automatically learn a motion histogram codebook and the set of ego-action categories. We quantitatively evaluate our approach on both in-house and public YouTube videos and demonstrate robust ego-action categorization across several sports genres. Comparative analysis shows that our approach outperforms other state-of-the-art topic models with respect to both classification accuracy and computational speed. Preliminary results indicate that on average, the categorical content of a 10 minute video sequence can be indexed in under 5 seconds.
Kris Makoto Kitani, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto
CVPR3
2011 Aesthetic quality classification of photographs based on color harmony
abstract
Aesthetic quality classification plays an important role in how people organize large photo collections. In particular, color harmony is a key factor in the various aspects that determine the perceived quality of a photo, and it should be taken into account to improve the performance of automatic aesthetic quality classification. However, the existing models of color harmony take only simple color patterns into consideration-e.g., patches consisting of a few colors-and thus cannot be used to assess photos with complicated color arrangements. In this work, we tackle the challenging problem of evaluating the color harmony of photos with a particular focus on aesthetic quality classification. A key point is that a photograph can be seen as a collection of local regions with color variations that are relatively simple. This led us to develop a method for assessing the aesthetic quality of a photo based on the photo's color harmony. We term the method `bags-of-color-patterns.' Results of experiments on a large photo collection with user-provided aesthetic quality scores show that our aesthetic quality classification method, which explicitly takes into account the color harmony of a photo, outperforms the existing methods. Results also show that the classification performance is improved by combining our color harmony feature with blur, edges, and saliency features that reflect the aesthetics of the photos.
Masashi Nishiyama, Takahiro Okabe, Imari Sato, Yoichi Sato 0001
CVPR4
2011 Inferring human gaze from appearance via adaptive linear regression
abstract
The problem of estimating human gaze from eye appearance is regarded as mapping high-dimensional features to low-dimensional target space. Conventional methods require densely obtained training samples on the eye appearance manifold, which results in a tedious calibration stage. In this paper, we introduce an adaptive linear regression (ALR) method for accurate mapping via sparsely collected training samples. The key idea is to adaptively find the subset of training samples where the test sample is most linearly representable. We solve the problem via l1-optimization and thoroughly study the key issues to seek for the best solution for regression. The proposed gaze estimation approach based on ALR is naturally sparse and low-dimensional, giving the ability to infer human gaze from variant resolution eye images using much fewer training samples than existing methods. Especially, the optimization procedure in ALR is extended to solve the subpixel alignment problem simultaneously for low resolution test eye images. Performance of the proposed method is evaluated by extensive experiments against various factors such as number of training samples, feature dimensionality and eye image resolution to verify its effectiveness.
Feng Lu 0005, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001
ICCV4
2011 Attention Prediction in Egocentric Video Using Motion and Visual Saliency
Kentaro Yamada, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto, Kazuo Hiraki
PSIVT (1)4
2011 Early facial expression recognition with high-frame rate 3D sensing
abstract
This work investigates a new challenging problem: how to exactly recognize facial expression as early as possible, while most works generally focus on improving the recognition rate of facial expression recognition. The features of facial expressions in their early stage are unfortunately very sensitive to noise due to their low intensity. So, we propose a novel wavelet spectral subtraction method to spatio-temporally refine the subtle facial expression features. Moreover, in order to achieve early facial expression recognition, we newly introduce an early AdaBoost algorithm for facial expression recognition problem. Experiments using our database established by using a high-frame rate 3D sensing showed that the proposed method has a promising performance on early facial expression recognition.
Lumei Su, Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato, Yoichi Sato 0001
SMC6
2010 Fast Spectral Reflectance Recovery Using DLP Projector
Shuai Han 0006, Imari Sato, Takahiro Okabe, Yoichi Sato 0001
ACCV (1)4
2010 Video Temporal Super-Resolution Based on Self-similarity
Mihoko Shimano, Takahiro Okabe, Imari Sato, Yoichi Sato 0001
ACCV (1)4
2010 Calibration-free gaze sensing using saliency maps
abstract
We propose a calibration-free gaze sensing method using visual saliency maps. Our goal is to construct a gaze estimator only using eye images captured from a person watching a video clip. The key is treating saliency maps of the video frames as probability distributions of gaze points. To efficiently identify gaze points from saliency maps, we aggregate saliency maps based on the similarity of eye appearances. We establish mapping between eye images to gaze points by Gaussian process regression. The experimental result shows that the proposed method works well with different people and video clips and achieves 6 degrees of accuracy, which is useful for estimating a person's attention on monitors.
Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001
CVPR3
2010 Recognition of Blurred Faces via Facial Deblurring Combined with Blur-Tolerant Descriptors
abstract
Blur is often present in real-world images and significantly affects the performance of face recognition systems. To improve the recognition of blurred faces, we propose a new approach which inherits the advantages of two recent methods. The idea consists of first reducing the amount of blur in the images via deblurring and then extracting blur-tolerant descriptors for recognition. We assess our analysis on real blurred face images (FRGC 1.0 database) and also on face images artificially degraded by focus blur (FERET database), demonstrating significant performance enhancement compared to the state-of-the-art.
Abdenour Hadid, Masashi Nishiyama, Yoichi Sato 0001
ICPR3
2010 Recovery of audio-to-video synchronization through analysis of cross-modality correlation
Yuyu Liu, Yoichi Sato 0001
Pattern Recognit. Lett.2
2010 Detecting Forgery From Static-Scene Video Based on Inconsistency in Noise Level Functions
abstract
Recently developed video editing techniques have enabled us to create realistic synthesized videos. Therefore, using video data as evidence in places such as courts of law requires a method to detect forged videos. In this study, we developed an approach to detect suspicious regions in a video of a static scene on the basis of the noise characteristics. The image signal contains irradiance-dependent noise the variance of which is described by a noise level function (NLF) as a function of irradiance. We introduce a probabilistic model providing the inference of an NLF that controls the characteristics of the noise at each pixel. Forged pixels in the regions clipped from another video camera can be differentiated by using maximum a posteriori estimation for the noise model when the NLFs of the regions are inconsistent with the rest of the video. We demonstrate the effectiveness of our proposed method by adapting it to videos recorded indoors and outdoors. The proposed method enables us to highly accurately evaluate the per-pixel authenticity of the given video, which achieves denser estimation than prior work based on block-level validation. In addition, the proposed method can be applied to various kinds of videos such as those contaminated by large noise and recorded with any scan formats, which limits the applicability of the existing methods.
Michihiro Kobayashi, Takahiro Okabe, Yoichi Sato 0001
IEEE Trans. Inf. Forensics Secur.3
2009 Image Enhancement of Low-Light Scenes with Near-Infrared Flash Images
Sosuke Matsui, Takahiro Okabe, Mihoko Shimano, Yoichi Sato 0001
ACCV (1)4
2009 Attached shadow coding: Estimating surface normals from shadows under unknown reflectance and lighting conditions
abstract
We present a novel technique, termed attached shadow coding, for estimating surface normals from shadows when the reflectance and lighting conditions are unknown. Our key idea is encoding surface points via attached shadows observed under different light source directions and then estimating surface normals on the basis of the similarity of the attached shadow codes. Because shadows do not rely on reflectance properties, our method is applicable to surfaces with various complex reflectances such as anisotropic and composite materials. Moreover, our method is robust against noise because it takes advantage of the combination of weak constraints imposed by a number of light sources. We theoretically show that the distance between the codes at two surface points is equal to the angle between the corresponding surface normals under the assumption of uniform lighting and a convex object. Our method embeds high-dimensional codes into a 3D surface normal space so that the inter-code distances are preserved. Furthermore, we extend the method in order to alleviate the effects of nonuniform lighting and cast shadows. Experimental results demonstrate the effectiveness of our method.
Takahiro Okabe, Imari Sato, Yoichi Sato 0001
ICCV3
2009 Using individuality to track individuals: Clustering individual trajectories in crowds using local appearance and frequency trait
abstract
In this work, we propose a method for tracking individuals in crowds. Our method is based on a trajectory-based clustering approach that groups trajectories of image features that belong to the same person. The key novelty of our method is to make use of a person's individuality, that is, the gait features and the temporal consistency of local appearance to track each individual in a crowd. Gait features in the frequency domain have been shown to be an effective biometric cue in discriminating between individuals, and our method uses such features for tracking people in crowds for the first time. Unlike existing trajectory-based tracking methods, our method evaluates the dissimilarity of trajectories with respect to a group of three adjacent trajectories. In this way, we incorporate the temporal consistency of local patch appearance to differentiate trajectories of multiple people moving in close proximity. Our experiments show that the use of gait features and the temporal consistency of local appearance contributes to significant performance improvement in tracking people in crowded scenes.
Daisuke Sugimura, Kris Makoto Kitani, Takahiro Okabe, Yoichi Sato 0001, Akihiro Sugimoto
ICCV4
2009 Visual localization of non-stationary sound sources
abstract
Sound source can be visually localized by analyzing the correlation between audio and visual data. To correctly analyze this correlation, the sound source is required to be stationary in a scene to date. We introduce a technique that localizes the non-stationary sound sources to overcome this limitation. The problem is formulated as finding the optimal visual trajectories that best represent the movement of the sound source over the pixels in a spatio-temporal volume. Using a beam search, we search these optimal visual trajectories by maximizing the correlation between the newly introduced audiovisual features of inconsistency. An incremental correlation evaluation with mutual information is developed here, which significantly reduces the computational cost. The correlations computed along the optimal trajectories are finally incorporated into a segmentation technique to localize a sound source region in the first visual frame of the current time window. Experimental results demonstrate the effectiveness of our method.
Yuyu Liu, Yoichi Sato 0001
ACM Multimedia2
2009 Sensation-based photo cropping
abstract
This paper proposes a novel method for automatically cropping a photo using a quality classifier that assesses whether the cropped region is agreeable to users. We statistically build this quality classifier using large photo collections available on websites where people manually insert quality scores to photos. We first trim the original image and then decide on the candidates for cropping. We find the cropped region with the highest quality score by applying the quality classifier to the candidates. Current automatic photo cropping techniques search for attention grabbing regions that consist of salient pixels from the original photo. They are not always pleasant to users because they do not take into account the quality of the cropped region. Our method with the quality classifier outperforms a state-of-the-art method that takes into consideration only the user's attention for automatic photo cropping.
Masashi Nishiyama, Takahiro Okabe, Yoichi Sato 0001, Imari Sato
ACM Multimedia3
2009 Detecting Video Forgeries Based on Noise Characteristics
Michihiro Kobayashi, Takahiro Okabe, Yoichi Sato 0001
PSIVT3
2009 Recognizing Multiple Objects via Regression Incorporating the Co-occurrence of Categories
Takahiro Okabe, Yuhi Kondo, Kris Makoto Kitani, Yoichi Sato 0001
PSIVT4
2009 Pose-Invariant Facial Expression Recognition Using Variable-Intensity Templates
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001
Int. J. Comput. Vis.5
2008 Combining Stochastic and Deterministic Search for Pose-Invariant Facial Expression Recognition
abstract
We propose a novel method for pose-invariant facial expression recognition from monocular video sequences that combines stochastic and determinis-tic search processes. We use the simple face model called variable-intensity template, which can be prepared with very little time and effort. We tackle the two issues found in previous work on the variable-intensity template: low accuracy in head pose estimation, and assumption violations due to external intensity changes such as illumination change. We mitigate these issues by introducing the deterministic approach into the stochastic approach imple-mented as a particle filter. Our experiment demonstrates significant improve-ments in recognition performance for horizontal and vertical head orienta-tions in the range of ±40 degrees and ±20 degrees, respectively, from the frontal view. 1
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001
BMVC5
2008 An Incremental Learning Method for Unconstrained Gaze Estimation
Yusuke Sugano, Yasuyuki Matsushita, Yoichi Sato 0001, Hideki Koike
ECCV (3)3
2008 Recovering audio-to-video synchronization by audiovisual correlation analysis
abstract
Audio-to-video synchronization (AV-sync) may drift and is difficult to recover without dedicated human effort. In this work, we develop an interactive method to recover the drifted AV-sync by audiovisual correlation analysis. Given a video segment, a user specifies a rough time span during which a person is speaking. Our system first detects a speaker region using face detection. It then does a two-stage search to find the optimum AV-drift that can maximize the average audiovisual correlation inside the speaker region. The correlation is evaluated using quadratic mutual information with kernel density estimation. AV-sync is finally recovered by the detected optimum AV-drift. Experimental results demonstrate the effectiveness of our method.
Yuyu Liu, Yoichi Sato 0001
ICPR2
2008 Recognizing Overlapped Human Activities from a Sequence of Primitive Actions via Deleted Interpolation
abstract
The high-level recognition of human activity requires a priori hierarchical domain knowledge as well as a means of reasoning based on that knowledge. Based on insights from perceptual psychology, the problem of human action recognition is approached on the understanding that activities are hierarchical, temporally constrained and at times temporally overlapped. A hierarchical Bayesian network (HBN) based on a stochastic context-free grammar (SCFG) is implemented to address the hierarchical nature of human activity recognition. Then it is shown how the HBN is applied to different substrings in a sequence of primitive action symbols via deleted interpolation (DI) to recognize temporally overlapped activities. Results from the analysis of action sequences based on video surveillance data show the validity of the approach.
Kris Makoto Kitani, Yoichi Sato 0001, Akihiro Sugimoto
Int. J. Pattern Recognit. Artif. Intell.2
2008 Recovering the Basic Structure of Human Activities from Noisy Video-Based Symbol Strings
abstract
In recent years stochastic context-free grammars have been shown to be effective in modeling human activities because of the hierarchical structures they represent. However, most of the research in this area has yet to address the issue of learning the activity grammars from a noisy input source, namely, video. In this paper, we present a framework for identifying noise and recovering the basic activity grammar from a noisy symbol string produced by video. We identify the noise symbols by finding the set of non-noise symbols that optimally compresses the training data, where the optimality of compression is measured using an MDL criterion. We show the robustness of our system to noise and its effectiveness in learning the basic structure of human activity, through experiments with artificial data and a real video sequence from a local convenience store.
Kris Makoto Kitani, Yoichi Sato 0001, Akihiro Sugimoto
Int. J. Pattern Recognit. Artif. Intell.2
2007 Pose-Invariant Facial Expression Recognition Using Variable-Intensity Templates
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001
ACCV (1)5
2007 Shape Reconstruction Based on Similarity in Radiance Changes under Varying Illumination
abstract
This paper presents a technique for determining an object's shape based on the similarity of radiance changes observed at points on its surface under varying illumination. To examine the similarity, we use an observation vector that represents a sequence of pixel intensities of a point on the surface under different lighting conditions. Assuming convex objects under distant illumination and orthographic projection, we show that the similarity between two observation vectors is closely related to the similarity between the surface normals of the corresponding points. This enables us to estimate the object's surface normals solely from the similarity of radiance changes under unknown distant lighting by using dimensionality reduction. Unlike most previous shape reconstruction methods, our technique neither assumes particular reflection models nor requires reference materials. This makes our method applicable to a wide variety of objects made of different materials.
Imari Sato, Takahiro Okabe, Qiong Yu, Yoichi Sato 0001
ICCV4
2007 Learning motion patterns and anomaly detection by Human trajectory analysis
abstract
In this paper, we propose a novel method to learn motion patterns and detect anomalies by human trajectory analysis. Human trajectories are various, for example, moving, roaming, pausing, and so on. But, current approaches for the analysis of motion patterns are effective only in understanding simple trajectories. We aim to understand complicated human trajectories with long-term observation. To deal with spatial and temporal features of trajectories, we employ HMM (Hidden Markov Model) to model time-series features of human positions. Next, a similarity matrix of HMM mutual distances is formed. MDS (Multi-Dimensional Scaling) based on eigenvector decomposition provides projected coordinates of trajectories in low-dimensional space. Then we apply k-means clustering to projected data in order to acquire human motion patterns. Anomalies can be detected by the use of likelihood scores for HMM representing motion patterns. We tested the proposed method by real-world trajectories data observed in a small store. Experimental result shows that our method accurately finds typical motion patterns and unusual trajectories.
Naohiko Suzuki, Kosuke Hirasawa, Kenichi Tanaka, Yoshinori Kobayashi, Yoichi Sato 0001, Yozo Fujino
SMC5
2007 Appearance Sampling of Real Objects for Variable Illumination
Imari Sato, Takahiro Okabe, Yoichi Sato 0001
Int. J. Comput. Vis.3
2006 Effects of Image Segmentation for Approximating Object Appearance Under Near Lighting
Takahiro Okabe, Yoichi Sato 0001
ACCV (1)2
2006 Face Recognition Under Varying Illumination Based on MAP Estimation Incorporating Correlation Between Surface Points
Mihoko Shimano, Kenji Nagao, Takahiro Okabe, Imari Sato, Yoichi Sato 0001
ACCV (1)5
2006 3D Head Tracking using the Particle Filter with Cascaded Classifiers
abstract
We propose a method for real-time people tracking using multiple cameras. The particle filter framework is known to be effective for tracking people, but most of existing methods adopt only simple perceptual cues such as color histogram or contour similarity for hypothesis evaluation. To improve the robustness and accuracy of tracking more sophisticated hypothesis evaluation is indispensable. We therefore present a novel technique for human head tracking using cascaded classifiers based on AdaBoost and Haar-like features for hypothesis evaluation. In addition, we use multiple classifiers, each of which is trained respectively to detect one direction of a human head. During real-time tracking the most suitable classifier is adaptively selected by considering each hypothesis and known camera position. Our experimental results demonstrate the effectiveness and robustness of our method. 1 1
Yoshinori Kobayashi, Daisuke Sugimura, Yoichi Sato 0001, Kousuke Hirasawa, Naohiko Suzuki, Hiroshi Kage, Akihiro Sugimoto
BMVC3
2006 Gaze Estimation from Low Resolution Images
Yasuhiro Ono, Takahiro Okabe, Yoichi Sato 0001
PSIVT3
2005 Using Extended Light Sources for Modeling Object Appearance under Varying Illumination
abstract
In this study, we demonstrate the effectiveness of using extended light sources for modeling the appearance of an object for varying illumination. Extended light sources have a radiance distribution that is similar to that of the Gaussian function and have the potential of functioning as a low-pass filter when the appearance of an object is sampled under them. This enables us to obtain a set of basis images of an object for variable illumination from input images of the object taken under those light sources without suffering aliasing caused by insufficient sampling of its appearance. Furthermore, extended light sources are useful in terms of reducing high contrast in image intensities due to specular and diffuse reflection components. This helps us observe both specular and diffuse reflection components of an object in the same image taken with a single shutter speed. We have tested our proposed approach based on extended light sources with objects of complex appearance that are generally difficult to model using image-based modeling techniques.
Imari Sato, Takahiro Okabe, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV3
2005 Motion Control of Self-Moving Trays for Human Supporting Production Cell "Attentive Workbench"
abstract
We propose “Attentive Workbench (AWB),” a new cell production system in which an intelligent system supports human workers. Using cameras, projectors, planar motor driven self-moving trays and other devices, the system supports workers from both physical and information aspects, recognizing worker’s condition and intention. This paper outlines AWB and deals with physical assembly support using self-moving parts trays. Two different schemes, centralized and decentralized, for controlling multiple parts trays are evaluated through experiments and simulations. The centralized control scheme is found to have the higher performance than the decentralized control scheme.
Masao Sugi, Makoto Nikaido, Yusuke Tamura, Jun Ota 0001, Tamio Arai, Kiyoshi Kotani, Kiyoshi Takamasu, Seiichi Shin, Hiromasa Suzuki, Yoichi Sato 0001
ICRA10
2005 Arrangement planning for multiple self-moving trays in human supporting production cell "attentive workbench"
abstract
We have proposed "attentive workbench (AWB)", a new cell production system in which an intelligent system supports human workers. Recognizing worker's condition and intention, the system supports workers from both physical and information aspects. This paper deals with physical assembly support using self-moving parts trays. The system delivers necessary assembly parts to workers and clears finished products quickly. On transporting large products, multiple trays are subjected to form a rigid body working as a large single tray. In order to realize this arrangement of trays in real-time, a planning method based on the priority scheme and heuristic rule is proposed. The present method is evaluated through simulations. A demonstration of assembly support using real self-moving trays is shown.
Makoto Nikaido, Masao Sugi, Yusuke Tamura, Jun Ota 0001, Tamio Arai, Kiyoshi Kotani, Kiyoshi Takamasu, Akio Yamamoto, Seiichi Shin, Hiromasa Suzuki, Yoichi Sato 0001
IROS11
2004 Spherical Harmonics vs. Haar Wavelets: Basis for Recovering Illumination from Cast Shadows
Takahiro Okabe, Imari Sato, Yoichi Sato 0001
CVPR (1)3
2004 Video Content Manipulation by Means of Content Annotation and Nonsymbolic Gestural Interfaces
Burin Anuchitkittikul, Masashi Okamoto, Sadao Kurohashi, Toyoaki Nishida, Yoichi Sato 0001
KES5
2003 Object Recognition Based on Photometric Alignment Using RANSAC
abstract
For object recognition under varying illumination conditions, we propose a method based on photometric alignment. The photometric alignment is known as a technique that models both diffuse reflection components and attached shadows under a distant point light source by using three basis images. However, in order to reliably reproduce these components in a test image, we have to take into account outliers such as specular reflection components and shadows in the test image. Accordingly, our proposed method utilizes Random Sample Consensus (RANSAC), which has been used successfully for estimating basis images. In the present study, we have conducted experiments using the Yale Face Database B and confirmed that a combination of the photometric alignment and RANSAC provides a simple but effective method for object recognition under varying illumination conditions.
Takahiro Okabe, Yoichi Sato 0001
CVPR (1)2
2003 Appearance Sampling for Obtaining A Set of Basis Images for Variable Illumination
abstract
Previous studies have demonstrated that the appearance of an object under varying illumination conditions can be represented by a low-dimensional linear subspace. A set of basis images spanning such a linear subspace can be obtained by applying the principal component analysis (PCA) for a large number of images taken under different lighting conditions. While the approaches based on PCA have been used successfully for object recognition under varying illumination conditions, little is known about how many images would be required in order to obtain the basis images correctly. In this study, we present a novel method for analytically obtaining a set of basis images of an object for arbitrary illumination from input images of the object taken under a point light source. The main contribution of our work is that we show that a set of lighting directions can be determined for sampling images of an object depending on the spectrum of the object's BRDF in the angular frequency domain such that a set of harmonic images can be obtained analytically based on the sampling theorem on spherical harmonics. In addition, unlike the previously proposed techniques based on spherical harmonics, our method does not require the 3D shape and reflectance properties of an object used for rendering harmonics images of the object synthetically.
Imari Sato, Takahiro Okabe, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV3
2003 Illumination from Shadows
abstract
In this paper, we introduce a method for recovering an illumination distribution of a scene from image brightness inside shadows cast by an object of known shape in the scene. In a natural illumination condition, a scene includes both direct and indirect illumination distributed in a complex way, and it is often difficult to recover an illumination distribution from image brightness observed on an object surface. The main reason for this difficulty is that there is usually not adequate variation in the image brightness observed on the object surface to reflect the subtle characteristics of the entire illumination. In this study, we demonstrate the effectiveness of using occluding information of incoming light in estimating an illumination distribution of a scene. Shadows in a real scene are caused by the occlusion of incoming light and, thus, analyzing the relationships between the image brightness and the occlusions of incoming light enables us to reliably estimate an illumination distribution of a scene even in a complex illumination environment. This study further concerns the following two issues that need to be addressed. First, the method combines the illumination analysis with an estimation of the reflectance properties of a shadow surface. This makes the method applicable to the case where reflectance properties of a surface are not known a priori and enlarges the variety of images applicable to the method. Second, we introduce an adaptive sampling framework for efficient estimation of illumination distribution. Using this framework, we are able to avoid a unnecessarily dense sampling of the illumination and can estimate the entire illumination distribution more efficiently with a smaller number of sampling directions of the illumination distribution. To demonstrate the effectiveness of the proposed method, we have successfully tested the proposed method by using sets of real images taken in natural illumination conditions with different surface materials of shadow regions.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
2002 Two-handed drawing on augmented desk system
abstract
This paper describes a two-handed drawing tool developed on our augmented desk system. Using our real-time finger tracking method, a user can draw and manipulate objects interactively by his/her own finger/hand. Based on the former work on two-handed interaction, different roles are assigned to each hand. The right hand is used to draw and to manipulate objects. Using gesture recognition, primitive objects can be drawn by users' handwriting. On the other hand, the left hand is used to manipulate menus and to assist the right hand. By closing all left hand fingers, users can initiate the appearance of structural radial menus around their left hands, and can select appropriate items by using a left hand finger. The left hand is also used to assist in the performance of drawing tasks, e.g., specifying the center of a circle or top-left corner of a rectangle, or specifying the object to be copied.
Xinlei Chen, Hideki Koike, Yasuto Nakanishi, Kenji Oka, Yoichi Sato 0001
AVI5
2002 Vision-Based Face Tracking System for Large Displays
Yasuto Nakanishi, Takashi Fujii, Kotaro Kitajima, Yoichi Sato 0001, Hideki Koike
UbiComp4
2001 Stability Issues in Recovering Illumination Distribution from Brightness in Shadows
abstract
The paper describes a robust method for estimating, in a reliable manner, the illumination distribution of a real scene from shadows in a given image. In general, shadows in a scene are caused by the occlusion of incoming light; image brightness inside shadows have great potential for providing distinct clues to the illumination distribution of the scene. Taking advantage of this fact, we recently proposed to estimate the illumination distribution of a real scene from a single image of the scene. The proposed method has been applied successfully to real images with complex illumination distributions. Nevertheless, it was found that sometimes the method failed to provide a correct estimate of illumination distribution. Those failures stem from the fact that the method does not take into account several factors regarding the stability of illumination estimation. The study analyzes how much information is obtainable from a given image about the illumination distribution of the scene. In particular, we carefully examine the source of instability of using shadows obtained from a single image for the estimation in several aspects: blocked view of shadows by the object; limited sampling resolution for image brightness inside shadows; and the appropriate light model to approximate the illumination distribution of the scene. Based on this analysis, we propose a method that guarantees to reliably estimate the illumination distribution of a scene, regardless of the type of input image.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
CVPR (2)2
2001 SnapLink: Interactive Object Registration and Recognition for Augmented Desk Interface
Takahiro Nishi, Yoichi Sato 0001, Hideki Koike
INTERACT2
2001 Real-Time Input of 3D Pose and Gestures of a User's Hand and Its Applications for HCI
abstract
Introduces a method for tracking a user's hand in 3D and recognizing the hand's gesture in real time without the use of any invasive devices attached to the hand. Our method uses multiple cameras for determining the position and orientation of a user's hand moving freely in a 3D space. In addition, the method identifies pre-determined gestures in a fast and robust manner by using a neural network which has been properly trained beforehand. This paper also describes results of user study of our proposed method and several types of applications, including 3D object handling for a desktop system and a 3D walkthrough for a large immersive display system.
Yoichi Sato 0001, Makiko Saito, Hideki Koike
VR1
2001 Eigen-Texture Method: Appearance Compression and Synthesis Based on a 3D Model
abstract
Image-based and model-based methods are two representative rendering methods for generating virtual images of objects from their real images. However, both methods still have several drawbacks when we attempt to apply them to mixed reality where we integrate virtual images with real background images. To overcome these difficulties, we propose a new method, which we refer to as the Eigen-Texture method. The proposed method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface generated from a sequence of range images. The Eigen-Texture method is an example of a view-dependent texturing approach which combines the advantages of image-based and model-based approaches. No reflectance analysis of the object surface is needed, while an accurate 3D geometric model facilitates integration with other scenes. The paper describes the method and reports on its implementation.
Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi
IEEE Trans. Pattern Anal. Mach. Intell.2
2001 Integrating paper and digital information on EnhancedDesk: a method for realtime finger tracking on an augmented desk system
abstract
This article describes a design and implementation of an augmented desk system, named EnhancedDesk, which smoothly integrates paper and digital information on a desk. The system provides users an intelligent environment that automatically retrieves and displays digital information corresponding to the real objects (e.g., books) on the desk by using computer vision. The system also provides users direct manipulation of digital information by using the users' own hands and fingers for more natural and more intuitive interaction. Based on the experiments with our first prototype system, some critical issues on augmented desk systems were identified when trying to pursue rapid and fine recognition of hands and fingers. To overcome these issues, we developed a novel method for realtime finger tracking on an augmented desk system by introducing a infrared camera, pattern matching with normalized correlation, and a pan-tilt camera. We then show an interface prototype on EnhancedDesk. It is an application to a computer-supported learning environment, named Interactive Textbook. The system shows how effective the integration of paper and digital information is and how natural and intuitive direct manipulation of digital information with users' hands and fingers is.
Hideki Koike, Yoichi Sato 0001, Yoshinori Kobayashi
ACM Trans. Comput. Hum. Interact.2
2000 Interactive textbook and interactive Venn diagram: natural and intuitive interfaces on augmented desk system
abstract
This paper describes two interface prototypes which we have developed on our augmented desk interface system, EnhancedDesk. The first application is Interactive Textbook, which is aimed at providing an effective learning environment. When a student opens a page which describes experiments or simulations, Interactive Textbook automatically retrieves digital contents from its database and projects them onto the desk. Interactive Textbook also allows the student hands-on ability to interact with the digital contents. The second application is the Interactive Venn Diagram, which is aimed at supporting effective information retrieval. Instead of keywords, the system uses real objects such as books or CDs as keys for retrieval. The system projects a circle around each book; data corresponding the book are then retrieved and projected inside the circle. By moving two or more circles so that the circles intersect each other, the user can compose a Venn diagram interactively on the desk. We also describe the new technologies introduced in EnhancedDesk which enable us to implement these applications.
Hideki Koike, Yoichi Sato 0001, Yoshinori Kobayashi, Hiroaki Tobita, Motoki Kobayashi
CHI2
2000 Fast Tracking of Hands and Fingertips in Infrared Images for Augmented Desk Interface
abstract
We introduce a fast and robust method for tracking positions of the centers and the fingertips of both right and left hands. Our method makes use of infrared camera images for reliable detection of a user's hands, and uses a template matching strategy for finding fingertips. This method is an essential part of our augmented desk interface in which a user can, with natural hand gestures, simultaneously manipulate both physical objects and electronically projected objects on a desk, e.g., a textbook and related WWW pages. Previous tracking methods which are typically based on color segmentation or background subtraction simply do not perform well in this type of application because an observed color of human skin and image backgrounds may change significantly due to protection of various objects onto a desk. In contrast, our proposed method was shown to be effective even in such a challenging situation through demonstration in our augmented desk interface. This paper describes the details of our tracking method as well as typical applications in our augmented desk interface.
Yoichi Sato 0001, Yoshinori Kobayashi, Hideki Koike
FG1
2000 Robust Localization for 3D Object Recognition Using Local EGI and 3D Template Matching with M-Estimators
abstract
A teleoperated system in a robot greatly reduces the demands on the human operator, although some human intervention is still required to perform such tasks as insulator recognition, positional adjustments of the robot, and guidance toward electric lines and insulators. In order to automate some of the robot's capabilities, we have developed a 3D object-localization method for the robot's positional adjustment. The method is designed to be insensitive to noise, outliers and occlusions while, at the same time, it has optimal run-time efficiency. The main contribution of our algorithm is the use of an objective function which is specified to reduce the effect of noise and outliers in the range image and a method for minimizing this function. The objective function is efficiently minimized by dynamically recomputing correspondences as the pose improves. Our algorithm is general enough to be applied not only to our dual-armed mobile robots but also to other teleoperation robots. This algorithm should greatly reduce the burden of operators when applied. This paper first describes our algorithm, and then presents a performance evaluation.
Kentaro Kawamura, Kiminori Hasegawa, Yasuyuki Someya, Yoichi Sato 0001, Katsushi Ikeuchi
ICRA4
2000 Appearance-based visual learning and object recognition with illumination invariance
Kohtaro Ohba, Yoichi Sato 0001, Katsushi Ikeuchi
Mach. Vis. Appl.2
1999 Eigen-Texture Method: Appearance Compression Based on 3D Model
abstract
Image-based and model-based methods are two representative rendering methods for generating virtual images of objects from their real images. Extensive research on these two methods has been made in CV and CG communities. However, both methods still have several drawbacks when it comes to applying them to the mixed reality where we integrate such virtual images with real background images. To overcome these difficulties, we propose a new method which we refer to as the Eigen-Texture method. The proposed method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface. The 3D model is generated from a sequence of range images. The Eigen-Texture method is practical because it does not require any detailed reflectance analysis of the object surface, and has great advantages due to the accurate 3D geometric models. This paper describes the method, and reports on its implementation.
Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi
CVPR2
1999 Measurement of Surface Orientations of Transparent Objects Using Polarization in Highlight
abstract
This paper proposes a method for obtaining surface orientations of transparent objects using polarization in highlight. Since the highlight, the specular component of reflection light from objects, is observed only near the specular direction, it appears merely on limited parts on an object surface. In order to obtain orientations of a whole object surface, we employ a spherical extended light source. This paper reports its experimental apparatus, a shape recovery algorithm, and its performance evaluation.
Megumi Saito, Hiroshi Kashiwagi, Yoichi Sato 0001, Katsushi Ikeuchi
CVPR3
1999 Illumination Distribution from Shadows
abstract
The image irradiance of a three-dimensional object is known to be the function of three components: the distribution of light sources, the shape, and reflectance of a real object surface. In the past, recovering the shape and reflectance of an object surface from the recorded image brightness has been intensively investigated. On the other hand, there has been little progress in recovering illumination from the knowledge of the shape and reflectance of a real object. In this paper, we propose a new method for estimating the illumination distribution of a real scene from image brightness observed on a real object surface in that scene. More specifically, we recover the illumination distribution of the scene from a radiance distribution inside shadows cast by an object of known shape onto another object surface of known shape and reflectance. By using the occlusion information of the incoming light, we are able to reliably estimate the illumination distribution of a real scene, even in a complex illumination environment.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
CVPR2
1999 Appearance Compression and Synthesis based on 3D Model for Mixed Reality
abstract
Rendering photorealistic virtual objects from their real images is one of the main research issues in mixed reality systems. We previously proposed the Eigen-Texture method (K. Nishino et al., 1999), a new rendering method for generating virtual images of objects from their real images to deal with the problems posed by past work in image based methods and model based methods. Eigen-Texture method samples appearances of a real object under various illumination and viewing conditions, and compresses them in the 2D coordinate system defined on the 3D model surface. However, we had a serious limitation in our system, due to the alignment problem of the 3D model and color images. We deal with this limitation by solving the alignment problem; we do this by using the method originally designed by P. Viola (1995). The paper describes the method and reports on how we implement it.
Ko Nishino, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV2
1999 Illumination Distribution from Brightness in Shadows: Adaptive Estimation of Illumination Distribution with Unknown Reflectance Properties in Shadow Regions
abstract
An approach to extract watersheds and watercourses, as well as their corresponding valleys and hills, from images with subpixel precision is proposed. The critical points of the terrain are essential as the starting points for the construction of these separatrices. They are extracted efficiently with subpixel precision using an approach based on derivatives of Gaussian filters. The separatrices are extracted by integrating their defining differential equation. Finally, the hills and valleys are constructed by an efficient graph search algorithm. Examples show the quality of the results that can be achieved with the proposed approach.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV2
1999 Acquiring a Radiance Distribution to Superimpose Virtual Objects onto a Real Scene
abstract
This paper describes a new method for superimposing virtual objects with correct shadings onto an image of a real scene. Unlike the previously proposed methods, our method can measure a radiance distribution of a real scene automatically and use it for superimposing virtual objects appropriately onto a real scene. First, a geometric model of the scene is constructed from a pair of omnidirectional images by using an omnidirectional stereo algorithm. Then, radiance of the scene is computed from a sequence of omnidirectional images taken with different shutter speeds and mapped onto the constructed geometric model. The radiance distribution mapped onto the geometric model is used for rendering virtual objects superimposed onto the scene image. As a result, even for a complex radiance distribution, our method can superimpose virtual objects with convincing shadings and shadows cast onto the real scene. We successfully tested the proposed method by using real images to show its effectiveness.
Imari Sato, Yoichi Sato 0001, Katsushi Ikeuchi
IEEE Trans. Vis. Comput. Graph.2
1998 Appearance Based Visual Learning and Object Recognition with Illumination Invariance
Kohtaro Ohba, Yoichi Sato 0001, Katsushi Ikeuchi
ACCV (2)2
1998 Measuring Object Surface Shape and Reflectance Properties
Yoichi Sato 0001, Mark D. Wheeler, Katsushi Ikeuchi
ACCV (2)1
1998 Consensus Surfaces for Modeling 3D Objects from Multiple Range Images
abstract
In this paper, we present a robust method for creating a triangulated surface mesh from multiple range images. Our method merges a set of range images into a volumetric implicit-surface representation which is converted to a surface mesh using a variant of the marching-cubes algorithm. Unlike previous techniques based on implicit-surface representations, our method estimates the signed distance to the object surface by finding a consensus of locally coherent observations of the surface. We call this method the consensus-surface algorithm. This algorithm effectively eliminates many of the troublesome effects of noise and extraneous surface observations without sacrificing the accuracy of the resulting surface. We utilize octrees to represent volumetric implicit surfaces-effectively reducing the computation and memory requirements of the volumetric representation without sacrificing accuracy of the resulting surface. We present results which demonstrate that our consensus-surface algorithm can construct accurate geometric models from rather noisy input range data.
Mark D. Wheeler, Yoichi Sato 0001, Katsushi Ikeuchi
ICCV2
1998 Localization of insulators in electric distribution systems by using 3D template matching from multiple range images
abstract
Kyushu Electric has developed a dual-armed mobile-robot for use in electricity distribution systems. Although some human intervention is still required, the robot greatly reduces the demands on the human operator. In order to automate some of the robot's capabilities, we have developed a 3D object-localization method for robot positional adjustment. The method is designed to be insensitive to noise and outliers while, at the same time, it has optimal run-time efficiency. The paper first describes our algorithm, and then presents a performance evaluation.
Kentaro Kawamura, Mark D. Wheeler, Osamu Yamashita, Yoichi Sato 0001, Katsushi Ikeuchi
IROS4
1997 Visual learning and object verification with illumination invariance
abstract
This paper describes a method for recognizing partially occluded objects to realize a bin-picking task under different levels of illumination brightness by using the eigenspace analysis. In the proposed method, a measured color in the RGB color space is transformed into the HSV color space. Then, the hue of the measured color, which is invariant to change in illumination brightness and direction, is used for recognizing multiple objects under different levels of illumination conditions. The proposed method was applied to real images of multiple objects under different illumination conditions, and the objects were recognized and localized successfully.
Kohtaro Ohba, Yoichi Sato 0001, Katsushi Ikeuchi
IROS2
1997 Object shape and reflectance modeling from observation
abstract
An object model for computer graphics applications should contain two aspects of information: shape and reflectance properties of the object.A number of techniques have been developed for modeling object shapes by observing real objects.In contrast, attempts to model reflectance properties of real objects have been rather limited.In most cases, modeled reflectance properties are too simple or too complicated to be used for synthesizing realistic images of the object.In this paper, we propose a new method for modeling object reflectance properties, as well as object shapes, by observing real objects.First, an object surface shape is reconstructed by merging multiple range images of the object.By using the reconstructed object shape and a sequence of color images of the object, parameters of a reflection model are estimated in a robust manner.The key point of the proposed method is that, first, the diffuse and specular reflection components are separated from the color image sequence, and then, reflectance parameters of each reflection component are estimated separately.This approach enables estimation of reflectance properties of real objects whose surfaces show specularity as well as diffusely reflected lights.The recovered object shape and reflectance properties are then used for synthesizing object images with realistic shading effects under arbitrary illumination conditions.
Yoichi Sato 0001, Mark D. Wheeler, Katsushi Ikeuchi
SIGGRAPH1
1997 3D shape and reflectance morphing
abstract
The paper describes a new method for 3D shape and reflectance morphing of two real 3D objects. Our morphing method consists of two components: shape and reflectance property measurement, and smooth interpolation of those measured properties. Unlike other morphing techniques, the proposed method can create intermediate images with correct shading such as highlights and shadows.
Yoichi Sato 0001, Imari Sato, Katsushi Ikeuchi
Shape Modeling International1
1996 Reflectance Analysis for 3D Computer Graphics Model Generation
Yoichi Sato 0001, Katsushi Ikeuchi
CVGIP Graph. Model. Image Process.1
1993 Temporal-color space analysis of reflection
abstract
A method to analyze a sequence of color images is proposed. A series of images is examined in a four-dimensional space, which is called the temporal-color space, the axes of which are the three color axes (RGB) and one temporal axis. The significance of the temporal-color space lies in its ability to represent the change of image color with time. A conventional color space analysis yields the histogram of the colors in an image only at an instance of time. Conceptually, the two reflection components from the dichromatic reflection model, the specular reflection component and the body reflection component, form two subspaces in the temporal-color space. These two components can be extracted by principal component analysis. Using this fact, real color images are analyzed, and the two reflection components are separated successfully.>
Yoichi Sato 0001, Katsushi Ikeuchi
CVPR1