EDBT 2026 Demo / reviewers in the wild / expert
Yutaka Satoh
dblp:57/1765
· DBLP profile ↗
43ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-0638-0855ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 8 since 2021Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Red, Thinking Bad: Color Bias in Vision Language Models
Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh |
ICPR (2) | 5 |
| 2026 | Diffusion Noise Optimization for Synthetic VLM TrainingabstractRecent advances in image generation models have enabled the production of high-quality images, making synthetic images a promising alternative to real images for dataset construction. However, a critical challenge remains in that the performance of Vision–Language Models (VLMs) tends to degrade as the proportion of synthetic images in a dataset increases in conventional approaches. To alleviate the challenge, we introduce a plug-and-play dataset construction framework that enhances text-to-image diffusion models by optimizing their initial noise. Our method treats the initial noise as a learnable parameter and iteratively updates it to maximize text–image alignment based on multiple embedding models without retraining the generator. Since the initial noise plays a crucial role in determining the quality of the synthetic image, its optimization enables the search for initial conditions that yield semantically faithful and realistic images. By improving FID and text–image alignment compared to conventional latent diffusion model (LDM)-based methods, our approach produces synthetic images better suited for training. When CLIP models were trained on such images, it achieved up to +5.09% higher Average R@1 in zero-shot retrieval, +2.88% higher Average top-1 accuracy in zero-shot classification, and +5.05% higher performance in linear-probing. These results demonstrate that initial noise optimization is an effective and scalable strategy for enabling robust VLM training with synthetic images. Ren Ohkubo, Rintaro Yanagi, Hirokatsu Kataoka, Yutaka Satoh |
WACV | 4 |
| 2024 | Formula-Supervised Visual-Geometric Pre-training
Ryosuke Yamada, Kensho Hara, Hirokatsu Kataoka, Koshi Makihara, Nakamasa Inoue, Rio Yokota, Yutaka Satoh |
ECCV (22) | 7 |
| 2024 | Subtle-Diff: A Dataset for Precise Recognition of Subtle Differences Among Visually Similar ObjectsabstractVisual inspection robots used in factories and outdoor environments require the ability to accurately recognize visual differences between similar objects and further verbalize the recognition results to present the differences to humans. Despite the application of Large Language Models (LLMs) and multimodal LLMs across various domains, our research highlights their insufficiency in verbalizing nuanced differences across images. To address this, we leveraged LLMs and image generation AI to develop a dataset aimed at assessing difference recognition capabilities. We introduced two novel tasks using this dataset: selecting images based on their visual differences and a conditional difference captioning task, and evaluated existing Vision-Language Models (VLMs) on these tasks. Our findings reveal that advanced models like GPT-4V can describe subtle differences with comparative expressions, yet they fall short of matching human performance across all attributes. This discrepancy between model and human recognition, especially in identifying easily discernible differences, suggests that most current models lack the ability to directly compare image pairs for difference detection. Consequently, we propose a new model that incorporates an image-text similarity approach in the difference recognition task, showing superior performance over existing models, including GPT-4V. Our dataset and findings will contribute to advancements in differencing objects and improve robotic applications in visual inspection and object picking. The dataset is available at DICTA challenge page. Fumiya Matsuzawa, Yue Qiu 0001, Yanjun Sun, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh |
IROS | 6 |
| 2023 | Question Generation for Uncertainty Elimination in Referring Expressions in 3D EnvironmentsabstractWe introduce a new task of question generation to eliminate the uncertainty of referring expressions in 3D indoor environments (3D-REQ). Referring to an object using natural language is one of the most common occurrences in daily human conversations; therefore, instructing robots to identify a certain object using natural language could be an essential task in var-ious robotic applications, such as room arrangement. However, human instructions are sometimes uncertain. Existing research on visual grounding using natural language in a 3D environment assumes that the referring expression can uniquely identify the object and does not consider that humans unconsciously give uncertain expressions. When faced with uncertainties, humans ask questions to gain further information. Inspired by the above observation, we propose a method that reduces uncertainty by asking questions when being given an obscure referring expression. The purpose of this method is to predict the positions of all candidate objects that satisfy the referring expressions in a 3D indoor environment and then to ask the appropriate questions to narrow down the target objects from them. To achieve this, we constructed a new 3D-REQ dataset, the input of which is a referring expression with uncertainties in the 3D environment and point clouds, and the output of which is the bounding boxes of all candidate objects satisfying the referring expression and a question to eliminate the uncertainty. To the best of our knowledge, 3D-REQ is the first effort to eliminate the uncertainty of referring expressions for object grounding in 3D environments. Fumiya Matsuzawa, Yue Qiu 0001, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh |
ICRA | 5 |
| 2023 | VirtualHome Action Genome: A Simulated Spatio-Temporal Scene Graph Dataset with Consistent Relationship LabelsabstractSpatio-temporal scene graph generation is an essential task in household activity recognition that aims to identify human-object interactions. Constructing a dataset with per-frame object region and consistent relationship annotations requires extremely high labor costs. Existing datasets sparsely annotate frames sampled from videos, resulting in the lack of dense spatio-temporal correlation in videos. Additionally, existing datasets contain inconsistent relationship annotations, leading to the problem of learning ambiguous temporal associations. Moreover, existing datasets mainly discuss relationships that can be inferred from a single frame, ignoring the significance of temporal associations. To resolve those issues, we created a simulated dataset with per-frame consistent annotations and introduced a range of relationships requiring both spatial and temporal context. Most existing methods explore spatial correlations within single images and do not explicitly consider the dynamic changes across frames. Therefore, we proposed a tracking-based approach that explicitly grasps spatio-temporal human-object interactions while simultaneously localizing humans and objects. Our proposed approach achieved state-of-the-art performance on scene graph generation and outperformed existing methods in scene graph localization by large margins on the proposed dataset. Moreover, the experiments show the efficacy of pre-training on the proposed dataset while adapting to a previous benchmark consisting of real daily videos, indicating the potential of the proposed dataset in real-world scenarios. Yue Qiu 0001, Yoshiki Nagasaki, Kensho Hara, Hirokatsu Kataoka, Ryota Suzuki 0006, Kenji Iwata, Yutaka Satoh |
WACV | 7 |
| 2023 | 3D Change Localization and Captioning from Dynamic Scans of Indoor ScenesabstractDaily indoor scenes often involve constant changes due to human activities. To recognize scene changes, existing change captioning methods focus on describing changes from two images of a scene. However, to accurately perceive and appropriately evaluate physical changes and then identify the geometry of changed objects, recognizing and localizing changes in 3D space is crucial. Therefore, we propose a task to explicitly localize changes in 3D bounding boxes from two point clouds and describe detailed scene changes, including change types, object attributes, and spatial locations. Moreover, we create a simulated dataset with various scenes, allowing generating data without labor costs. We further propose a framework that allows different 3D object detectors to be incorporated in the change detection process, after which captions are generated based on the correlations of different change regions. The proposed framework achieves promising results in both change detection and captioning. Furthermore, we also evaluated on data collected from real scenes. The experiments show that pretraining on the proposed dataset increases the change detection accuracy by +12.8% (mAP0.25) when applied to real-world data. We believe that our proposed dataset and discussion could provide both a new benchmark and in-sights for future studies in scene change understanding. Yue Qiu 0001, Shintaro Yamamoto, Ryosuke Yamada, Ryota Suzuki 0006, Hirokatsu Kataoka, Kenji Iwata, Yutaka Satoh |
WACV | 7 |
| 2022 | Can Vision Transformers Learn without Natural Images?abstractIs it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated labels, the recent use of natural images has resulted in problems related to privacy violation, inadequate fairness protection, and the need for labor-intensive annotations. In this paper, we experimentally verify that the results of formula-driven supervised learning (FDSL) framework are comparable with, and can even partially outperform, sophisticated self-supervised learning (SSL) methods like SimCLRv2 and MoCov2 without using any natural images in the pre-training phase. We also consider ways to reorganize FractalDB generation based on our tentative conclusion that there is room for configuration improvements in the iterated function system (IFS) parameter settings of such databases. Moreover, we show that while ViTs pre-trained without natural images produce visualizations that are somewhat different from ImageNet pre-trained ViTs, they can still interpret natural image datasets to a large extent. Finally, in experiments using the CIFAR-10 dataset, we show that our model achieved a performance rate of 97.8, which is comparable to the rate of 97.4 achieved with SimCLRv2 and 98.0 achieved with ImageNet. Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata, Nakamasa Inoue, Yutaka Satoh |
AAAI | 6 |
| 2022 | Pre-Training Without Natural ImagesabstractAbstract Is it possible to use convolutional neural networks pre-trained without any natural images to assist natural image understanding? The paper proposes a novel concept, Formula-driven Supervised Learning (FDSL). We automatically generate image patterns and their category labels by assigning fractals, which are based on a natural law. Theoretically, the use of automatically generated images instead of natural images in the pre-training phase allows us to generate an infinitely large dataset of labeled images. The proposed framework is similar yet different from Self-Supervised Learning because the FDSL framework enables the creation of image patterns based on any mathematical formulas in addition to self-generated labels. Further, unlike pre-training with a synthetic image dataset, a dataset under the framework of FDSL is not required to define object categories, surface texture, lighting conditions, and camera viewpoint. In the experimental section, we find a better dataset configuration through an exploratory study, e.g., increase of #category/#instance, patch rendering, image coloring, and training epoch. Although models pre-trained with the proposed Fractal DataBase (FractalDB), a database without natural images, do not necessarily outperform models pre-trained with human annotated datasets in all settings, we are able to partially surpass the accuracy of ImageNet/Places pre-trained models. The FractalDB pre-trained CNN also outperforms other pre-trained models on auto-generated datasets based on FDSL such as Bezier curves and Perlin noise. This is reasonable since natural objects and scenes existing around us are constructed according to fractal geometry. Image representation with the proposed FractalDB captures a unique feature in the visualization of convolutional layers and attentions. Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, Yutaka Satoh |
Int. J. Comput. Vis. | 8 |
| 2022 | Predicting Appearance of Vehicles From Blind Spots Based on Pedestrian Behaviors at CrossroadsabstractConventional prediction approaches for traffic scenes primarily predict the future states of visible objects (i.e., not in blind spots) based on their current observations. This study focused on the prediction of future states of objects in blind spots (e.g., those outside the filed-of-view or occluded regions) based on the current observations of other visible objects. We proposed a method that predicts the appearance of vehicles from a blind spot based on the behaviors of visible pedestrians who observe vehicles in the blind spot. Our proposed method utilizes a spatiotemporal 3D convolutional neural network and learns pedestrian behaviors for predictions. The method explicitly represents subtle motions and the surrounding environments of pedestrians using pose estimation and semantic segmentation. To conduct evaluation experiments, we built two datasets of videos capturing real traffic scenes. The datasets are collected by cameras with and without ego-motions. Using the datasets, we conducted experiments not only on simpler configurations but also on realistic traffic environments. Based on the experimental results, the following conclusions could be obtained: (i) our proposed method achieved a high performance at a level similar to that of humans in our prediction task, and predicted the appearance of vehicles from blind spots more than 1.5 s before they actually appeared. (ii) Explicit representations of pose and semantic masks captured information complementary to RGB videos, and ensembling the representations improved the prediction performance. (iii) Fine-tuning the models using videos with ego-motions is important to achieve good prediction in the videos captured by driving cars. Kensho Hara, Hirokatsu Kataoka, Masaki Inaba, Kenichi Narioka, Ryusuke Hotta, Yutaka Satoh |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Describing and Localizing Multiple Changes with TransformersabstractChange captioning tasks aim to detect changes in image pairs observed before and after a scene change and generate a natural language description of the changes. Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-change captioning dataset; (ii) We benchmark existing state-of-the-art methods of single change captioning on multi-change captioning; (iii) We further propose Multi-Change Captioning transformers (MCCFormers) that identify change regions by densely correlating different regions in image pairs and dynamically determines the related change regions with words in sentences. The proposed method obtained the highest scores on four conventional change captioning evaluation metrics for multi-change captioning. Additionally, our proposed method can separate attention maps for each change and performs well with respect to change localization. Moreover, the proposed framework outperformed the previous state-of-the-art methods on an existing change captioning benchmark, CLEVR-Change, by a large margin (+6.1 on BLEU-4 and +9.7 on CIDEr scores), indicating its general ability in change captioning tasks. The code and dataset are available at the project page1. Yue Qiu 0001, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki 0006, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh |
ICCV | 7 |
| 2021 | Viewpoint-agnostic Image RenderingabstractRendering an any-viewpoint image is extremely difficult for Generative Adversarial Networks. This is because conventional GANs do not understand 3D information under-lying a given viewpoint image such as an object shape and relationship between viewpoint and objects in 3D space. In this paper, we present how to perform a Viewpoint-Agnostic Image Rendering (VAIR), equipping a conditional GAN with a mechanism to reconstruct 3D information of the input view. VAIR realizes any-viewpoint image generation by manipulating a viewpoint in 3D space where the reconstructed instance shape is arranged. In addition, we convert the reconstructed 3D shape into a 2D representation for image-based conditional GAN, while preserving detail 3D information. The representation consists of a depth image and 2D semantic keypoint images, which are obtained by rendering the shape from a viewpoint. In the experiment, we evaluate using a CUB-200-2011 dataset, which contains few-samples biased a viewpoint such that covers only part of the target appearance. As a result, our VAIR clearly renders an any-viewpoint image. Hiroaki Aizawa, Hirokatsu Kataoka, Yutaka Satoh, Kunihito Kato |
WACV | 3 |
| 2020 | Pre-training Without Natural Images
Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, Yutaka Satoh |
ACCV (6) | 8 |
| 2020 | Disentangle, Assemble, and Synthesize: Unsupervised Learning to Disentangle Appearance and LocationabstractThe next step for the generative adversarial networks (GAN) is to learn representations that allow us to control only a certain factor in the image explicitly. Since such a representation of the factor is independent of other factors, the controllability obtained from these representations leads to interpretability by identifying the variation of the synthesized image and the transferability for downstream tasks by inference. However, since it is difficult to identify and strictly define latent factors, the annotation is laborious. Moreover, learning such representations by a GAN is challenging due to the complex generation process. Therefore, we resolve this limitation using a novel generative model that can disentangle latent space into the appearance, the x-axis, and the y-axis of the object, and reassemble these components in an unsupervised manner. Specifically, based on the concept of packing the appearance and location in each position of the feature map, we introduce a novel structural constraint technique that prevents these representations from interacting with each other. The proposed structural constraint promotes the disentanglement of these factors. In experiments, we found that the proposed method is simple but effective for controllability and allows us to control the appearance and location via latent space without supervision, as compared with the conditional GAN. Hiroaki Aizawa, Hirokatsu Kataoka, Yutaka Satoh, Kunihito Kato |
ICPR | 3 |
| 2020 | Joint Pedestrian Detection and Risk-level Prediction with Motion-Representation-by-DetectionabstractThe paper presents a pedestrian near-miss detector with temporal analysis that provides both pedestrian detection and risk-level predictions which are demonstrated on a self-collected database. Our work makes three primary contributions: (i) The framework of pedestrian near-miss detection is proposed by providing both a pedestrian detection and risk-level assignment. Specifically, we have created a Pedestrian Near-Miss (PNM) dataset that categorizes traffic near-miss incidents based on their risk levels (high-, low-, and no-risk). Unlike existing databases, our dataset also includes manually localized pedestrian labels as well as a large number of incident-related videos. (ii) Single-Shot MultiBox Detector with Motion Representation (SSD-MR) is implemented to effectively extract motion-based features in a detected pedestrian. (iii) Using the self-collected PNM dataset and SSD-MR, our proposed method achieved +19.38% (on risk-level prediction) and +13.00% (on joint pedestrian detection and risk-level prediction) higher scores than that of the baseline SSD and LSTM. Additionally, the running time of our system is over 50 fps on a graphics processing unit (GPU). Hirokatsu Kataoka, Teppei Suzuki, Kodai Nakashima, Yutaka Satoh, Yoshimitsu Aoki |
ICRA | 4 |
| 2019 | Incorporating 3D Information Into Visual Question AnsweringabstractWe propose a tactic of advancing Visual Question Answering (VQA) task by incorporating 3D information via multi-view images. Conventional VQA approaches, which reply an answer in words against a linguistic question about a given RGB image, have less ability to recognize geometrical information so that they tend to fail to count things or guess positional relationship. Moreover, they have no ability to determine blinded space, so it is not feasible to invent VQA function to robots which will work in highly-occluded real-world environments. To achieve the situation, we introduce a new multi-view VQA dataset along with an approach that incorporating 3D scene information directly captured from multi-view images into VQA without using depth images or employing SLAM. Our proposed approach achieves strong performance with an overall accuracy of 95.4% on the challenging multi-view VQA dataset setup, which contains relatively severe occlusion. This work also demonstrates the promising aspects of bridging the gap between 3D vision and language. Yue Qiu 0001, Yutaka Satoh, Ryota Suzuki 0006, Hirokatsu Kataoka |
3DV | 2 |
| 2019 | Unsupervised Out-of-context Action UnderstandingabstractThe paper presents an unsupervised out-of-context action (O2CA) paradigm that is based on facilitating understanding by separately presenting both human action and context within a video sequence. As a means of generating an unsupervised label, we comprehensively evaluate responses from action-based (ActionNet) and context-based (ContextNet) convolutional neural networks (CNNs). Additionally, we have created three synthetic databases based on the human action (UCF101, HMDB51) and motion capture (mocap) (SURREAL) datasets. We then conducted experimental comparisons between our approach and conventional approaches. We also compared our unsupervised learning method with supervised learning using an O2CA ground truth given by synthetic data. From the results obtained, we achieved a 96.8 score on Synth-UCF, a 96.8 score on Synth-HMDB, and 89.0 on SURREAL-O2CA with F-score. Hirokatsu Kataoka, Yutaka Satoh |
ICRA | 2 |
| 2019 | Foreground detection based on co-occurrence background model with hypothesis on degradation modification in dynamic scenes
Shun'ichi Kaneko, Manabu Hashimoto, Yutaka Satoh, Dong Liang 0008 |
Signal Process. | 4 |
| 2018 | Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?abstractThe purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels. Recently, the performance levels of 3D CNNs in the field of action recognition have improved significantly. However, to date, conventional research has only explored relatively shallow 3D architectures. We examine the architectures of various 3D CNNs from relatively shallow to very deep ones on current video datasets. Based on the results of those experiments, the following conclusions could be obtained: (i) ResNet-18 training resulted in significant overfitting for UCF-101, HMDB-51, and ActivityNet but not for Kinetics. (ii) The Kinetics dataset has sufficient data for training of deep 3D CNNs, and enables training of up to 152 ResNets layers, interestingly similar to 2D ResNets on ImageNet. ResNeXt-101 achieved 78.4% average accuracy on the Kinetics test set. (iii) Kinetics pretrained simple 3D architectures outperforms complex 2D architectures, and the pretrained ResNeXt-101 achieved 94.5% and 70.2% on UCF-101 and HMDB-51, respectively. The use of 2D CNNs trained on ImageNet has produced significant progress in various tasks in image. We believe that using deep 3D CNNs together with Kinetics will retrace the successful history of 2D CNNs and ImageNet, and stimulate advances in computer vision for videos. The codes and pretrained models used in this study are publicly available1. Kensho Hara, Hirokatsu Kataoka, Yutaka Satoh |
CVPR | 3 |
| 2018 | Anticipating Traffic Accidents With Adaptive Loss and Large-Scale Incident DBabstractIn this paper, we propose a novel approach for traffic accident anticipation through (i) Adaptive Loss for Early Anticipation (AdaLEA) and (ii) a large-scale self-annotated incident database for anticipation. The proposed AdaLEA allows a model to gradually learn an earlier anticipation as training progresses. The loss function adaptively assigns penalty weights depending on how early the model can anticipate a traffic accident at each epoch. Additionally, we construct a Near-miss Incident DataBase for anticipation. This database contains an enormous number of traffic near-miss incident videos and annotations for detail evaluation of two tasks, risk anticipation and risk-factor anticipation. In our experimental results, we found our proposal achieved the highest scores for risk anticipation (+6.6% better on mean average precision (mAP) and 2.36 sec earlier than previous work on the average time-to-collision (ATTC)) and risk-factor anticipation (+4.3% better on mAP and 0.70 sec earlier than previous work on ATTC). Hirokatsu Kataoka, Yoshimitsu Aoki, Yutaka Satoh |
CVPR | 4 |
| 2018 | Fashion Culture Database: Construction of Database for World-wide Fashion AnalysisabstractThe paper presents a novel concept that analyzes and visualizes worldwide fashion styles. Our goal is to reveal web-based viral fashion styles. To achieve the fashion-based analysis, we have collected fashion culture database (FCDB), which consists of 76 million geo-tagged images in 16 cosmopolitan cities. The database allows us to grasp a trend of mixed fashion styles with a fashion-based descriptor and codeword vector. In order to unveil web-based fashion trends in the FCDB, we applied a simple technique that is a temporal subtraction between consecutive codeword vectors in two different times. In the experiments, we show the analysis of fashion trends and fashion-based city similarity in a social media. As the result of large-scale data collection, we achieved world-level fashion visualization. Kaori Abe, Munetaka Minoguchi, Teppei Suzuki, Naofumi Akimoto, Yue Qiu 0001, Ryota Suzuki 0006, Kenji Iwata, Yutaka Satoh, Hirokatsu Kataoka |
ICARCV | 9 |
| 2018 | Semantic Change DetectionabstractChange detection is the study of detecting changes between two different images of a scene taken at different times. The change detection methodology can provide us information in which area images changed time by time. However, for application use, especially on disaster investigation, it is highly required to understand not only where but also what changes are occurred in high precision and resolution. The paper proposes the concept of semantic change detection, which involves intuitively inserting semantic meaning into detected change areas. We mainly focus on the novel semantic segmentation in addition to a conventional change detection approach. In order to solve this problem and obtain a high-level of performance, we propose an improvement to the hypercolumns representation, hereafter known as hypermaps, which effectively uses convolutional maps obtained from convolutional neural networks (CNNs). We also employ multi-scale feature representation captured by different image patches. We applied our method to the TSUNAMI panoramic change detection dataset (TSUNAMI dataset), and re-annotated the changed areas of the dataset via semantic classes. The results show that our multi-scale hypermaps provided outstanding performance on the re-annotated TSUNAMI dataset. Munetaka Minoguchi, Ryota Suzuki 0006, Akio Nakamura, Kenji Iwata, Yutaka Satoh, Hirokatsu Kataoka |
ICARCV | 6 |
| 2018 | Towards Good Practice for Action Recognition with Spatiotemporal 3D ConvolutionsabstractThe purpose of this study is to explore good practice for training convolutional neural networks (CNNs) with spatiotemporal three-dimensional (3D) kernels. Recently, 3D CNNs in the field of action recognition are rapidly developed, and the performance levels of them have improved significantly. However, to date, conventional research has mainly focused on their architecture, and has not sufficiently explored their training configurations. We conduct various experiments with different training configurations on Kinetics, UCF-101, and HMDB-51 datasets to share the knowledge of 3D CNNs for the research community. According to the results of those experiments, the following conclusions could be obtained. (i) Data augmentation by spatiotemporal random cropping improved the performance levels. (ii) Data augmentation by multi-scale spatial cropping increased the accuracies in most cases whereas multi-scale temporal cropping decreased them. (iii) A corner cropping strategy, which is previously shown as a good method for two-stream 2D CNNs, resulted lower accuracies for 3D CNNs compared with simple random cropping. (iv) Freezing early layers of 3D CNNs improved the performance levels when fine-tuning 3D CNNs on a relatively small dataset. Kensho Hara, Hirokatsu Kataoka, Yutaka Satoh |
ICPR | 3 |
| 2018 | Occlusion Handling Human Detection with Refocused ImagesabstractThe paper presents a novel robust human detection method based on camera array system to broaden the application range for human detection. Currently, even by using a deep neural network (DNN), it is difficult to detect a hardly occluded human. In the camera array system, we consider how to distinctly show a human occluded by an environmental condition. The generated refocused images by the camera array system allow us to remove the effect of the noises. Although refocused images have not been utilized in conventional human detection, we believe that the refocused images are beneficial for improving the detection performance, especially in severe conditions. To execute the experiments, we have collected Refocused Human DataBase (RHDB) with the camera array system. By using HOG+SVM with a monocular camera (at an almost random rate of 54.8%), the refocused images made the +10.1% improvement (64.9%) by noticeably showing a human. The combined representation of refocused images and AlexNet achieved 94.6% on the RHDB. Moreover, our final model recorded 98.0% with an attention-layer and fine-tuned parameters. Hirokatsu Kataoka, Shuhei Ohki, Kenji Iwata, Yutaka Satoh |
ICPR | 4 |
| 2018 | A Co-occurrence Background Model with Hypothesis on Degradation Modification for Object Detection in Strong Background ChangesabstractObject detection has become an indispensable part of video processing and current background models are sensitive to background changes. In this paper, we propose a novel background model using an algorithm called Co-occurrence Pixel-block Pairs (CPB) against background changes, such as illumination changes and background motion. We utilize the co-occurrence “pixel to block” structure to extract the spatial-temporal information of each pixel to build background model, and then employ an efficient evaluation strategy to identify the current state of each pixel, which is named as correlation dependent decision function. Furthermore, we also introduce a Hypothesis on Degradation Modification (HoD) into CPB structure to reinforce the robustness of CPB. Experimental results obtained from the dataset of the PETS 2001, AIST-Indoor, SBMnet and CDW-2012 databases show that our models can detect objects robustly in strong background changes. Shun'ichi Kaneko, Manabu Hashimoto, Yutaka Satoh, Dong Liang 0008 |
ICPR | 4 |
| 2018 | Drive Video Analysis for the Detection of Traffic Near-Miss IncidentsabstractBecause of their recent introduction, self-driving cars and advanced driver assistance system (ADAS) equipped vehicles have had little opportunity to learn, the dangerous traffic (including near-miss incident) scenarios that provide normal drivers with strong motivation to drive safely. Accordingly, as a means of providing learning depth, this paper presents a novel traffic database that contains information on a large number of traffic near-miss incidents that were obtained by mounting driving recorders in more than 100 taxis over the course of a decade. The study makes the following two main contributions: (i) In order to assist automated systems in detecting near-miss incidents based on database instances, we created a large-scale traffic near-miss incident database (NIDB) that consists of video clip of dangerous events captured by monocular driving recorders. (ii) To illustrate the applicability of NIDB traffic near-miss incidents, we provide two primary database-related improvements: parameter fine-tuning using various near-miss scenes from NIDB, and foreground/background separation into motion representation. Then, using our new database in conjunction with a monocular driving recorder, we developed a near-miss recognition method that provides automated systems with a performance level that is comparable to a human-level understanding of near-miss incidents (64.5% vs. 68.4% at near-miss recognition, 61.3% vs. 78.7% at near-miss detection). Hirokatsu Kataoka, Teppei Suzuki, Shoko Oikawa, Yasuhiro Matsui, Yutaka Satoh |
ICRA | 5 |
| 2017 | Illuminant-Camera Communication to Observe Moving Objects under Strong External Light by Spread Spectrum ModulationabstractMany algorithms of computer vision use light sources to illuminate objects to actively create situation appropriate to extract their characteristics. For example, the shape and reflectance are measured by a projector-camera system, and some human-machine or VR systems use projectors and displays for interaction. As existing active lighting systems usually assume no severe external lights to observe projected lights clearly, it is one of the limitations of active illumination. In this paper, we propose a method of energy-efficient active illumination in an environment with severe external lights. The proposed method extracts the light signals of illuminants by removing external light using spread spectrum modulation. Because an image sequence is needed to observe modulated signals, the proposed method extends signal processing to realize signal detection projected onto moving objects by combining spread spectrum modulation and spatio-temporal filtering. In the experiments, we apply the proposed method to a structured-light system under sunlight, to photometric stereo with external lights, and to insensible image embedding. Ryusuke Sagawa, Yutaka Satoh |
CVPR | 2 |
| 2017 | Text Detection in Traffic Informatory Signs Using Synthetic DataabstractTraffic informatory signs, which is a category of traffic signs and text-based signs, is very important to both drivers and intelligent transport systems. Previous studies have usually sought to extract text lines in signs to apply to optical character recognition (OCR) system, but they do not work well in real-world conditions with severe disturbance. In this paper, we report on our study of place name text detection and recognition on traffic informatory signs using convolutional neural networks (CNNs) and transform traditional text detection and recognition into word-level multi-class robust image classification. In our study, each place name corresponds to one class. Because the number of word classes is large and collecting real images for training dataset is difficult, we generate several synthetic datasets mainly by means of two methods and use them to train the CNN respectively. One method generates with standard templates of the signs, while the other method renders text in natural images which have no relationship with the signs. Our experimental results show that our method can achieve high levels of accuracy when reading traffic informatory signs in real-world conditions. Accuracy of 0.891 and 0.981 are achieved in the former and latter method of generating dataset, which verify that proposed methods are effective in our task. We also analyze the dependence of color channels in text detection during our task, which help generate the more efficacious synthetic dataset. Fangge Chen, Hirokatsu Kataoka, Yutaka Satoh |
ICDAR | 3 |
| 2016 | Recognition of Transitional Action for Short-Term Action Prediction using Discriminative Temporal CNN Feature
Hirokatsu Kataoka, Yudai Miyashita, Masaki Hayashi, Kenji Iwata, Yutaka Satoh |
BMVC | 5 |
| 2014 | Extended Co-occurrence HOG with Dense Trajectories for Fine-Grained Activity Recognition
Hirokatsu Kataoka, Kiyoshi Hashimoto, Kenji Iwata, Yutaka Satoh, Nassir Navab, Slobodan Ilic, Yoshimitsu Aoki |
ACCV (5) | 4 |
| 2013 | Co-occurrence-based adaptive background model for robust object detectionabstractAn illumination-invariant background model for detecting objects in dynamic scenes is proposed. It is robust in the cases of sudden illumination fluctuation as well as burst moving background. Unlike previous works, it distinguishes objects from a dynamic background using co-occurrence character between a target pixel and its supporting pixels in the form of multiple pixel pairs. Experiments used several challenging datasets that proved the robust performance of object detection in various environments. Dong Liang 0008, Shun'ichi Kaneko, Manabu Hashimoto, Kenji Iwata, Xinyue Zhao, Yutaka Satoh |
AVSS | 6 |
| 2013 | Robust feature descriptor and vehicle motion model with tracking-by-detection for active safetyabstractThe percentage of pedestrian deaths in traffic accidents is on the rise in Japan. In recent years, there have been calls for measures to be introduced to protect vulnerable road users such as pedestrians and cyclists. In this study, a method to detect and track pedestrians using an in-vehicle camera is presented to perform braking controls, warn the driver, and develop improved safety systems for pedestrians. We improved the technology of detecting pedestrians using highly accurate images obtained with a monocular camera. We were able to predict pedestrian activity by monitoring the images, and developed an algorithm with which to recognize pedestrians and their movements more accurately. The effectiveness of the algorithm was tested using images taken on real roads. For the feature descriptor, we used an extended co-occurrence histogram of oriented gradients (ECoHOG) that accumulated the integration of gradient intensities. In the tracking step, we applied an effective motion model using optical flow and the proposed feature descriptor ECoHOG in a tracking-by-detection framework. These techniques were verified using images captured on the real road. Hirokatsu Kataoka, Kimimasa Tamura, Yoshimitsu Aoki, Yasuhiro Matsui, Kenji Iwata, Yutaka Satoh |
IECON | 6 |
| 2013 | Robust face recognition using the GAP feature
Xinyue Zhao, Zaixing He, Shuyou Zhang 0001, Shun'ichi Kaneko, Yutaka Satoh |
Pattern Recognit. | 5 |
| 2011 | Robust adapted object detection under complex environmentabstractIn this paper, we present a novel robust technique for background subtraction in different complex conditions (e.g. sudden illumination changes, swaying leaves, and camera vibrations). Unlike the previous works, the proposed method utilizes multiple point pairs that exhibit a stable statistical intensity relationship as a background model. The intensity difference between pixels of the pair is much more stable than the intensity of a single pixel, especially in varying environments. Furthermore, our proposed method focuses more on the history of global spatial correlations between pixels than on the history of any given pixel or local spatial correlations. we also adopt an adapted judgement criterion to ensure our method displays well in real-time detection. The approach has been compared with the state of the art on videos from several challenging datasets (PETS, Wallflower, and i-Lids), demonstrating that superior object detection is achieved. Xinyue Zhao, Yutaka Satoh, Hidenori Takauji, Shun'ichi Kaneko, Kenji Iwata, Ryushi Ozaki |
AVSS | 2 |
| 2011 | Object detection based on a robust and accurate statistical multi-point-pair model
Xinyue Zhao, Yutaka Satoh, Hidenori Takauji, Shun'ichi Kaneko, Kenji Iwata, Ryushi Ozaki |
Pattern Recognit. | 2 |
| 2007 | Application of the Unusual Motion Detection Using CHLAC to the Video Surveillance
Kenji Iwata, Yutaka Satoh, Takumi Kobayashi 0001, Ikushi Yoda, Nobuyuki Otsu |
ICONIP (2) | 2 |
| 2006 | Hybrid Camera Surveillance System by Using Stereo Omni-directional System and Robust Human Detection
Kenji Iwata, Yutaka Satoh, Ikushi Yoda, Katsuhiko Sakaue |
PSIVT | 2 |
| 2006 | Moving object detection by mobile Stereo Omni-directional System (SOS) using spherical depth image
Sanae Shimizu, Kazuhiko Yamamoto, Caihua Wang, Yutaka Satoh, Hideki Tanahashi, Yoshinori Niwa |
Pattern Anal. Appl. | 4 |
| 2003 | Slant estimation for active vision using edge directions in omnidirectional imagesabstractIn this paper, we propose a novel method to estimate the slant of omnidirectional image sensors which can acquire images of an environment with a FOV close to 360/spl deg/ /spl times/ 180/spl deg/. The proposed method consists of two steps: a low resolution voting followed by a least squares method (LSM). First, the direction of each edge pixel in the omnidirectional image is computed and projected to the plane of Z = 1. The biggest peak formed by the vertical edges, which are prevalent in indoor and urban scenes, is detected. Then, edge directions whose projections to the voting plane pass near the peak are used to estimate the slant of the sensor by LSM. Experimental results in a real environment show the effectiveness of the proposed method. Caihua Wang, Hideki Tanahashi, Yutaka Satoh, Hidekazu Hirayu, Yutaka Sato, Yoshinori Niwa, Kazuhiko Yamamoto |
ICIP (1) | 3 |
| 2003 | Robust object detection for intelligent surveillance systems based on radial reach correlation (RRC)abstractThis paper describes a novel algorithm for robust object detection and segmentation, which is based on a new robust dissimilarity measure called as radial reach correlation (RRC). The capability of detecting moving objects from a complex background is one of the most fundamental technology for intelligent surveillance systems. The RRC is a new robust dissimilarity measure and has a well-formed probabilistic model of binary or normal density. The RRC evaluates the local texture between a background image and the current scene and realize robust object detection under poor conditions. To demonstrate the effectiveness of our approach, we present experimental results from real world all-directional images provided by the stereo omni-directional system (SOS). Yutaka Satoh, Caihua Wang, Yoshinori Niwa, Hideki Tanahashi, Kazuhiko Yamamoto |
IROS | 1 |
| 2003 | Event Detection for a Visual Surveillance System Using Stereo Omni-directional System
Hideki Tanahashi, Yutaka Satoh, Yoshinori Niwa, Kazuhiko Yamamoto |
KES | 3 |
| 2003 | Using selective correlation coefficient for robust image registration
Shun'ichi Kaneko, Yutaka Satoh, Satoru Igarashi |
Pattern Recognit. | 2 |
| 2000 | Non-restricted measurement of walking distanceabstractVertical walking distance is discussed. The horizontal distance is estimated by using the three dimensional acceleration of the subject's toe. Considering the movement of the foot, it is assumed that only the pitch angle of the foot changes and that the velocity of the foot is never negative. A three dimensional accelerometer is fixed to the subject's toe and measures the acceleration during the swing phase of the foot. A piezoelectric gyro is used to estimate the angle of the foot and to calculate its horizontal acceleration. The horizontal distance is obtained by integrating the horizontal acceleration twice each step. The vertical distance is calculated by integrating the change of atmospheric pressure during ascent or descent. A band pass filter is applied to reject natural changes of the atmospheric pressure and sensor noise. The integration of the output of the filter produces a value corresponding to the vertical distance. Experiments were performed in a horizontal corridor and on some stairs. The results show that the error of the horizontal distance estimated is less than 5.3% and 0.0% on average, and the vertical distance is less than 11.1%. Koichi Sagawa, Hikaru Inooka, Yutaka Satoh |
SMC | 3 |