VLDB 2026 Research / reviewers in the wild / expert
Hirokatsu Kataoka
dblp:128/7522
· DBLP profile ↗
69ranked-venue papers
13as first author
42since 2021 · last 2026
0000-0001-8844-165XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 7 first-author · 34 since 2021Artificial intelligence and machine learning · 46 · 9 first-author · 28 since 2021Systems, architecture and hardware · 12 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Red, Thinking Bad: Color Bias in Vision Language Models
Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh |
ICPR (2) | 4 |
| 2026 | Diffusion Noise Optimization for Synthetic VLM TrainingabstractRecent advances in image generation models have enabled the production of high-quality images, making synthetic images a promising alternative to real images for dataset construction. However, a critical challenge remains in that the performance of Vision–Language Models (VLMs) tends to degrade as the proportion of synthetic images in a dataset increases in conventional approaches. To alleviate the challenge, we introduce a plug-and-play dataset construction framework that enhances text-to-image diffusion models by optimizing their initial noise. Our method treats the initial noise as a learnable parameter and iteratively updates it to maximize text–image alignment based on multiple embedding models without retraining the generator. Since the initial noise plays a crucial role in determining the quality of the synthetic image, its optimization enables the search for initial conditions that yield semantically faithful and realistic images. By improving FID and text–image alignment compared to conventional latent diffusion model (LDM)-based methods, our approach produces synthetic images better suited for training. When CLIP models were trained on such images, it achieved up to +5.09% higher Average R@1 in zero-shot retrieval, +2.88% higher Average top-1 accuracy in zero-shot classification, and +5.05% higher performance in linear-probing. These results demonstrate that initial noise optimization is an effective and scalable strategy for enabling robust VLM training with synthetic images. Ren Ohkubo, Rintaro Yanagi, Hirokatsu Kataoka, Yutaka Satoh |
WACV | 3 |
| 2025 | Formula-Supervised Sound Event Detection: Pre-Training Without Real DataabstractIn this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effectiveness for sound event detection (SED). The SED task, which involves estimating the types and timings of sound events, is particularly challenged by the difficulty of acquiring a sufficient quantity of accurately labeled training data. Moreover, it is well known that manually annotated labels often contain noises and are significantly influenced by the subjective judgment of annotators. To address these challenges, we propose a novel pretraining method that utilizes a synthetic dataset, Formula-SED, where acoustic data are generated solely based on mathematical formulas. The proposed method enables large-scale pre-training by using the synthesis parameters applied at each time step as ground truth labels, thereby eliminating label noise and bias. We demonstrate that large-scale pre-training with Formula-SED significantly enhances model accuracy and accelerates training, as evidenced by our results in the DESED dataset used for DCASE2023 Challenge Task 4. The project page is at https://yutoshibata07.github.io/Formula-SED/. Yuto Shibata, Keitaro Tanaka, Yoshiaki Bando, Keisuke Imoto, Hirokatsu Kataoka, Yoshimitsu Aoki |
ICASSP | 5 |
| 2025 | AgroBench: Vision-Language Model Benchmark in Agriculture
Risa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka, Masaki Onishi, Yoshitaka Ushiku |
ICCV | 3 |
| 2025 | AnimalClue: Recognizing Animals by their Traces
Risa Shinoda, Nakamasa Inoue, Iro Laina, Christian Rupprecht 0001, Hirokatsu Kataoka |
ICCV | 5 |
| 2025 | Viewpoint-Dependent 3D Visual Grounding for Mobile Robotsabstract3D visual grounding is the task of identifying objects in spatial environments based on textual descriptions, enabling natural language interactions between humans and robots. However, existing studies overlook viewpoint-dependent texts expressions, such as "the chair to your right", despite their frequent use in human instructions. In this paper, we introduce a novel problem setting focused on viewpoint-dependent texts and present a new dataset that incorporates the robot’s viewpoint. We conducted three experiments to analyze the dataset’s difficulty and characteristics by comparing existing models and newly designed models that take viewpoint as an additional input. Our results indicate that considering viewpoint is crucial for the object selection process in our task. In addition, we also found that the difficulty of the task varies depending on how the object is described in the text. Shogo Iwakata, Ryosuke Oshima, Hideki Tsunashima, Hirokatsu Kataoka, Shigeo Morishima |
ICIP | 5 |
| 2025 | Can Masking Background and Object Reduce Static Bias for Zero-Shot Action Recognition?
Takumi Fukuzawa, Kensho Hara, Hirokatsu Kataoka, Toru Tamaki |
MMM (4) | 3 |
| 2025 | Approximate Domain Unlearning for Vision-Language ModelsabstractPre-trained Vision-Language Models (VLMs) exhibit strong generalization capabilities, enabling them to recognize a wide range of objects across diverse domains without additional training. However, they often retain irrelevant information beyond the requirements of specific target downstream tasks, raising concerns about computational efficiency and potential information leakage. This has motivated growing interest in approximate unlearning, which aims to selectively remove unnecessary knowledge while preserving overall model performance. Existing approaches to approximate unlearning have primarily focused on {\em class unlearning}, where a VLM is retrained to fail to recognize specified object classes while maintaining accuracy for others. However, merely forgetting object classes is often insufficient in practical applications. For instance, an autonomous driving system should accurately recognize {\em real} cars, while avoiding misrecognition of {\em illustrated} cars depicted in roadside advertisements as {\em real} cars, which could be hazardous. In this paper, we introduce {\em Approximate Domain Unlearning (ADU)}, a novel problem setting that requires reducing recognition accuracy for images from specified domains (e.g., {\em illustration}) while preserving accuracy for other domains (e.g., {\em real}). ADU presents new technical challenges: due to the strong domain generalization capability of pre-trained VLMs, domain distributions are highly entangled in the feature space, making naive approaches based on penalizing target domains ineffective. To tackle this limitation, we propose a novel approach that explicitly disentangles domain distributions and adaptively captures instance-specific domain information. Extensive experiments on four multi-domain benchmark datasets demonstrate that our approach significantly outperforms strong baselines built upon state-of-the-art VLM tuning techniques, paving the way for practical and fine-grained unlearning in VLMs. Code : https://kodaikawamura.github.io/Domain_Unlearning/. Kodai Kawamura, Yuta Goto, Rintaro Yanagi, Hirokatsu Kataoka, Go Irie |
NeurIPS | 4 |
| 2024 | Real-SRGD: Enhancing Real-World Image Super-Resolution with Classifier-Free Guided Diffusion
Kenji Doi, Shuntaro Okada, Ryota Yoshihashi, Hirokatsu Kataoka |
ACCV (5) | 4 |
| 2024 | Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation
Ryota Yoshihashi, Yuya Otsuka, Kenji Doi, Tomohiro Tanaka, Hirokatsu Kataoka |
ACCV (5) | 5 |
| 2024 | Watermark-embedded Adversarial Examples for Copyright Protection against Diffusion ModelsabstractDiffusion Models (DMs) have shown remarkable capa-bilities in various image-generation tasks. However, there are growing concerns that DMs could be used to imitate unauthorized creations and thus raise copyright issues. To address this issue, we propose a novel framework that em-beds personal watermarks in the generation of adversarial examples. Such examples can force DMs to generate images with visible watermarks and prevent DMs from imitating unauthorized images. We construct a generator based on conditional adversarial networks and design three losses (adversarial loss, GAN loss, and perturbation loss) to gen-erate adversarial examples that have subtle perturbation but can effectively attack DMs to prevent copyright violations. Training a generator for a personal watermark by our method only requires 5–10 samples within 2–3 minutes, and once the generator is trained, it can generate adver-sarial examples with that watermark significantly fast (0.2s per image). We conduct extensive experiments in various conditional image-generation scenarios. Compared to ex-isting methods that generate images with chaotic textures, our method adds visible watermarks on the generated images, which is a more straightforward way to indicate copy-right violations. We also observe that our adversarial exam-ples exhibit good transferability across unknown generative models. Therefore, this work provides a simple yet powerful way to protect copyright from DM-based imitation. Peifei Zhu, Tsubasa Takahashi 0001, Hirokatsu Kataoka |
CVPR | 3 |
| 2024 | Scaling Backwards: Minimal Synthetic Pre-Training?
Ryu Tadokoro, Ryosuke Yamada, Yuki Markus Asano, Iro Laina, Christian Rupprecht 0001, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka |
ECCV (15) | 9 |
| 2024 | Rethinking Image Super-Resolution from Training Data Perspectives
Go Ohtani, Ryu Tadokoro, Ryosuke Yamada, Yuki Markus Asano, Iro Laina, Christian Rupprecht 0001, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka, Yoshimitsu Aoki |
ECCV (17) | 9 |
| 2024 | Formula-Supervised Visual-Geometric Pre-training
Ryosuke Yamada, Kensho Hara, Hirokatsu Kataoka, Koshi Makihara, Nakamasa Inoue, Rio Yokota, Yutaka Satoh |
ECCV (22) | 3 |
| 2024 | Pseudo-Outlier Synthesis Using Q-Gaussian Distributions for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection, which aims to determine whether an input is outside the training data distribution or not, is an indispensable task in many computer vision applications. In many of the previous studies on OOD detection for image classification, the class-conditional distribution of visual features is assumed to be a Gaussian. However, this may not be a reasonable assumption because unseen outliers do not always follow a Gaussian distribution. In this study, we investigated the potential effects of non-Gaussian distributions by using an OOD detection method based on Tsallis statistics, in which the family of q-Gaussian distributions involving short- and long-tail distributions are used for synthesizing pseudo outlier features for improving the effectiveness of training. In experiments on six image classification datasets, we show that the proposed method achieves good results in the comparison method. In addition, we find that samples with a smaller hem than the Gaussian distribution by all datasets by ablation studies of the tail of the distribution improve the performance of OOD detection. Ryu Tadokoro, Eisuke Yamagata, Yusuke Kondo, Kensho Hara, Hirokatsu Kataoka, Nakamasa Inoue |
ICASSP | 6 |
| 2024 | On the Relationship Between Double Descent of CNNs and Shape/Texture Bias Under Learning Process
Shun Iwase, Shuya Takahashi, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka, Eisaku Maeda |
ICPR (25) | 6 |
| 2024 | Guided by the Way: The Role of On-the-route Objects and Scene Text in Enhancing Outdoor NavigationabstractIn outdoor environments, Vision-and-Language Navigation (VLN) requires an agent to rely on multi-modal cues from real-world urban environments and natural language instructions. While existing outdoor VLN models predict actions using a combination of panorama and instruction features, this approach ignores objects in the environment and learns data bias to fail navigation. According to our preliminary findings, most instances of navigation failure in previous models were due to turning or stopping at the wrong place. In contrast, humans intuitively frequently use identifiable objects or store names as reference landmarks, ensuring accurate turns and stops, especially in unfamiliar places. To address this insight gap, we propose an Object-Attention VLN (OAVLN) model that helps the agent focus on relevant objects during training and understand the environment better. Our model outperforms previous methods in all evaluation metrics under both seen and unseen scenarios on two existing benchmark datasets, Touchdown and map2seq. Yanjun Sun, Yue Qiu 0001, Yoshimitsu Aoki, Hirokatsu Kataoka |
ICRA | 4 |
| 2024 | Subtle-Diff: A Dataset for Precise Recognition of Subtle Differences Among Visually Similar ObjectsabstractVisual inspection robots used in factories and outdoor environments require the ability to accurately recognize visual differences between similar objects and further verbalize the recognition results to present the differences to humans. Despite the application of Large Language Models (LLMs) and multimodal LLMs across various domains, our research highlights their insufficiency in verbalizing nuanced differences across images. To address this, we leveraged LLMs and image generation AI to develop a dataset aimed at assessing difference recognition capabilities. We introduced two novel tasks using this dataset: selecting images based on their visual differences and a conditional difference captioning task, and evaluated existing Vision-Language Models (VLMs) on these tasks. Our findings reveal that advanced models like GPT-4V can describe subtle differences with comparative expressions, yet they fall short of matching human performance across all attributes. This discrepancy between model and human recognition, especially in identifying easily discernible differences, suggests that most current models lack the ability to directly compare image pairs for difference detection. Consequently, we propose a new model that incorporates an image-text similarity approach in the difference recognition task, showing superior performance over existing models, including GPT-4V. Our dataset and findings will contribute to advancements in differencing objects and improve robotic applications in visual inspection and object picking. The dataset is available at DICTA challenge page. Fumiya Matsuzawa, Yue Qiu 0001, Yanjun Sun, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh |
IROS | 5 |
| 2024 | Learnable Cube-based Video Encryption for Privacy-Preserving Action RecognitionabstractWith the development of cloud services and machine learning, there has been an inevitable need to enhance privacy and security when serving video recognition models. Although existing image encryption methods can be used to address this issue, applying them frame by frame to videos is insufficient in two respects: model performance degradation and security strength. In this paper, we propose a novel encryption approach for privacy-preserving action recognition. It consists of two encrypting operations; Learnable Cube-based Video Encryption (LCVE) and ViT Scrambling. LCVE is video encryption based on spatio-temporal cubes, which has a large key space and can provide robust privacy protection. ViT Scrambling encrypts the Vision Transformer (ViT) model, which enables it to recognize the encrypted videos in the same manner as unencrypted videos without modifying the model architecture or fine-tuning on the encrypted data. We evaluate our method in an action recognition task with seven datasets containing a variety of action classes as well as motion and visual patterns. Empirical results demonstrate that LCVE combined with ViT Scrambling can preserve video privacy while recognizing action in encrypted videos as well as unencrypted videos. As a result, our approach outperforms existing privacy-preserving action recognition methods. Yuchi Ishikawa, Masayoshi Kondo, Hirokatsu Kataoka |
WACV | 3 |
| 2023 | Primitive Geometry Segment Pre-training for 3D Medical Image Segmentation
Ryu Tadokoro, Ryosuke Yamada, Kodai Nakashima, Hirokatsu Kataoka |
BMVC | 5 |
| 2023 | Graph Representation for Order-aware Visual TransformationabstractThis paper proposes a new visual reasoning formulation that aims at discovering changes between image pairs and their temporal orders. Recognizing scene dynamics and their chronological orders is a fundamental aspect of human cognition. The aforementioned abilities make it possible to follow step-by-step instructions, reason about and analyze events, recognize abnormal dynamics, and restore scenes to their previous states. However, it remains unclear how well current AI systems perform in these capabilities. Although a series of studies have focused on identifying and describing changes from image pairs, they mainly consider those changes that occur synchronously, thus neglecting potential orders within those changes. To address the above issue, we first propose a visual transformation graph structure for conveying order-aware changes. Then, we benchmarked previous methods on our newly generated dataset and identified the issues of existing methods for change order recognition. Finally, we show a significant improvement in order-aware change recognition by introducing a new model that explicitly associates different changes and then identifies changes and their orders in a graph representation. Yue Qiu 0001, Yanjun Sun, Fumiya Matsuzawa, Kenji Iwata, Hirokatsu Kataoka |
CVPR | 5 |
| 2023 | Visual Atoms: Pre-Training Vision Transformers with Sinusoidal WavesabstractFormula-driven supervised learning (FDSL) has been shown to be an effective method for pre-training vision transformers, where ExFractalDB-21k was shown to exceed the pre-training effect of ImageNet-21k. These studies also indicate that contours mattered more than textures when pre-training vision transformers. However, the lack of a systematic investigation as to why these contour-oriented synthetic datasets can achieve the same accuracy as real datasets leaves much room for skepticism. In the present work, we develop a novel methodology based on circular harmonics for systematically investigating the design space of contour-oriented synthetic datasets. This allows us to efficiently search the optimal range of FDSL parameters and maximize the variety of synthetic images in the dataset, which we found to be a critical factor. When the resulting new dataset VisualAtom-21k is used for pre-training ViT-Base, the top-1 accuracy reached 83.7% when fine-tuning on ImageNet-1k. This is only 0.5% difference from the top-1 accuracy (84.2%) achieved by the JFT-300M pre-training, even though the scale of images is 1/14. Unlike JFT-300M which is a static dataset, the quality of synthetic datasets will continue to improve, and the current work is a testament to this possibility. FDSL is also free of the common issues associated with real images, e.g. privacy/copyright issues, labeling costs/errors, and ethical biases. Sora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hirokatsu Kataoka, Rio Yokota |
CVPR | 4 |
| 2023 | Pre-training Vision Transformers with Very Limited Synthesized ImagesabstractFormula-driven supervised learning (FDSL) is a pre-training method that relies on synthetic images generated from mathematical formulae such as fractals. Prior work on FDSL has shown that pre-training vision transformers on such synthetic datasets can yield competitive accuracy on a wide range of downstream tasks. These synthetic images are categorized according to the parameters in the mathematical formula that generate them. In the present work, we hypothesize that the process for generating different instances for the same category in FDSL, can be viewed as a form of data augmentation. We validate this hypothesis by replacing the instances with data augmentation, which means we only need a single image per category. Our experiments shows that this one-instance fractal database (OFDB) performs better than the original dataset where instances were explicitly generated. We further scale up OFDB to 21,000 categories and show that it matches, or even surpasses, the model pre-trained on ImageNet-21k in ImageNet-1k fine-tuning. The number of images in OFDB is 21k, whereas ImageNet-21k has 14M. This opens new possibilities for pre-training vision transformers with much smaller datasets. Hirokatsu Kataoka, Sora Takashima, Edgar Josafat Martinez-Noriega, Rio Yokota, Nakamasa Inoue |
ICCV | 2 |
| 2023 | SegRCDB: Semantic Segmentation via Formula-Driven Supervised LearningabstractPre-training is a strong strategy for enhancing visual models to efficiently train them with a limited number of labeled images. In semantic segmentation, creating annotation masks requires an intensive amount of labor and time, and therefore, a large-scale pre-training dataset with semantic labels is quite difficult to construct. Moreover, what matters in semantic segmentation pre-training has not been fully investigated. In this paper, we propose the Segmentation Radial Contour DataBase (SegRCDB), which for the first time applies formula-driven supervised learning for semantic segmentation. SegRCDB enables pre-training for semantic segmentation without real images or any manual semantic labels. SegRCDB is based on insights about what is important in pre-training for semantic segmentation and allows efficient pre-training. Pre-training with SegRCDB achieved higher mIoU than the pre-training with COCO-Stuff for fine-tuning on ADE-20k and Cityscapes with the same number of training images. SegRCDB has a high potential to contribute to semantic segmentation pre-training and investigation by enabling the creation of large datasets without manual annotation. The SegRCDB dataset will be released under a license that allows research and commercial use. Code is available at: https://github.com/dahlian00/SegRCDB Risa Shinoda, Ryo Hayamizu, Kodai Nakashima, Nakamasa Inoue, Rio Yokota, Hirokatsu Kataoka |
ICCV | 6 |
| 2023 | Frequency-aware GAN for Adversarial Manipulation GenerationabstractImage manipulation techniques have drawn growing concerns as manipulated images might cause morality and security problems. Various methods have been proposed to detect manipulations and achieved promising performance. However, these methods might be vulnerable to adversarial attacks. In this work, we design an Adversarial Manipulation Generation (AMG) task to explore the vulnerability of image manipulation detectors. We first propose an optimal loss function and extend existing attacks to generate adversarial examples. We observe that existing spatial attacks cause large degradation in image quality and find the loss of high-frequency detailed components might be its major reason. Inspired by this observation, we propose a novel adversarial attack that incorporates both spatial and frequency features into the GAN architecture to generate adversarial examples. We further design an encoder-decoder architecture with skip connections of high-frequency components to preserve fine details. We evaluated our method on three image manipulation detectors (FCN, ManTra-Net and MVSS-Net) with three benchmark datasets (DEFACTO, CASIAv2 and COVER). Experiments show that our method generates adversarial examples significantly fast (0.01s per image), preserves better image quality (PSNR 30% higher than spatial attacks), and achieves a high attack success rate. We also observe that the examples generated by AMG can fool both classification and segmentation models, which indicates better transferability among different tasks. Peifei Zhu, Genki Osada, Hirokatsu Kataoka, Tsubasa Takahashi 0001 |
ICCV | 3 |
| 2023 | Scapegoat Generation for Privacy Protection from DeepfakeabstractTo protect privacy and prevent malicious use of deepfake, current studies propose methods that interfere with the generation process, such as detection and destruction approaches. However, these methods suffer from sub-optimal generalization performance to unseen models and add undesirable noise to the original image. To address these problems, we propose a new problem formulation for deepfake prevention: generating a "scapegoat image" by modifying the style of the original input in a way that is recognizable as an avatar by the user, but impossible to reconstruct the real face. Even in the case of malicious deepfake, the privacy of the users is still protected. To achieve this, we introduce an optimization-based editing method that utilizes GAN inversion to discourage deepfake models from generating similar scapegoats. We validate the effectiveness of our proposed method through quantitative and user studies. Gido Kato, Yoshihiro Fukuhara, Mariko Isogawa, Hideki Tsunashima, Hirokatsu Kataoka, Shigeo Morishima |
ICIP | 5 |
| 2023 | Question Generation for Uncertainty Elimination in Referring Expressions in 3D EnvironmentsabstractWe introduce a new task of question generation to eliminate the uncertainty of referring expressions in 3D indoor environments (3D-REQ). Referring to an object using natural language is one of the most common occurrences in daily human conversations; therefore, instructing robots to identify a certain object using natural language could be an essential task in var-ious robotic applications, such as room arrangement. However, human instructions are sometimes uncertain. Existing research on visual grounding using natural language in a 3D environment assumes that the referring expression can uniquely identify the object and does not consider that humans unconsciously give uncertain expressions. When faced with uncertainties, humans ask questions to gain further information. Inspired by the above observation, we propose a method that reduces uncertainty by asking questions when being given an obscure referring expression. The purpose of this method is to predict the positions of all candidate objects that satisfy the referring expressions in a 3D indoor environment and then to ask the appropriate questions to narrow down the target objects from them. To achieve this, we constructed a new 3D-REQ dataset, the input of which is a referring expression with uncertainties in the 3D environment and point clouds, and the output of which is the bounding boxes of all candidate objects satisfying the referring expression and a question to eliminate the uncertainty. To the best of our knowledge, 3D-REQ is the first effort to eliminate the uncertainty of referring expressions for object grounding in 3D environments. Fumiya Matsuzawa, Yue Qiu 0001, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh |
ICRA | 4 |
| 2023 | Traffic Incident Database with Multiple Labels Including Various Perspective Environmental InformationabstractTraffic accident recognition is essential in developing automated driving and Advanced Driving Assistant System technologies. A large dataset of annotated traffic accidents is necessary to improve the accuracy of traffic accident recognition using deep learning models. Conventional traffic accident datasets provide annotations on the presence or absence of traffic accidents and other teacher labels, improving traffic accident recognition performance. However, the labels annotated in conventional datasets need to be more comprehensive to de-scribe traffic accidents in detail. Therefore, we propose V-TIDB, a large-scale traffic accident recognition dataset annotated with various environmental information as multi-labels. Our proposed dataset aims to improve the performance of traffic accident recognition by annotating ten types of environmental information as teacher labels in addition to the presence or absence of traffic accidents. V-TIDB is constructed by collecting many videos from the Internet and annotating them with appropriate environmental information. In our experiments, we compare the performance of traffic accident recognition when only labels related to the presence or absence of traffic accidents are trained and when environmental information is added as a multi-label. In the second experiment, we compare the performance of the training with only “contact level,” which represents the severity of the traffic accident, and the performance with environmental information added as a multi-label. The results showed that 6 out of 10 environmental information labels improved the performance of recognizing the presence or absence of traffic accidents. In the experiment on the degree of recognition of traffic accidents, the performance of recognition of car wrecks and contacts was improved for all environmental information. These experiments show that V-TIDB can be used to learn traffic accident recognition models that take environmental information into account in detail and can be used for appropriate traffic accident analysis. Shota Nishiyama, Takuma Saito, Go Ohtani, Hirokatsu Kataoka, Kensho Hara |
IROS | 5 |
| 2023 | Diffusion-based Holistic Texture Rectification and SynthesisabstractWe present a novel framework for rectifying occlusions and distortions in degraded texture samples from natural images. Traditional texture synthesis approaches focus on generating textures from pristine samples, which necessitate meticulous preparation by humans and are often unattainable in most natural images. These challenges stem from the frequent occlusions and distortions of texture samples in natural images due to obstructions and variations in object surface geometry. To address these issues, we propose a framework that synthesizes holistic textures from degraded samples in natural images, extending the applicability of exemplar-based texture synthesis techniques. Our framework utilizes a conditional Latent Diffusion Model (LDM) with a novel occlusion-aware latent transformer. This latent transformer not only effectively encodes texture features from partially-observed samples necessary for the generation process of the LDM, but also explicitly captures long-range dependencies in samples with large occlusions. To train our model, we introduce a method for generating synthetic data by applying geometric transformations and free-form mask generation to clean textures. Experimental results demonstrate that our framework significantly outperforms existing methods both quantitatively and quantitatively. Furthermore, we conduct comprehensive ablation studies to validate the different components of our proposed framework. Results are corroborated by a perceptual user study which highlights the efficiency of our proposed approach. Guoqing Hao, Satoshi Iizuka, Kensho Hara, Edgar Simo-Serra, Hirokatsu Kataoka, Kazuhiro Fukui |
SIGGRAPH Asia | 5 |
| 2023 | VirtualHome Action Genome: A Simulated Spatio-Temporal Scene Graph Dataset with Consistent Relationship LabelsabstractSpatio-temporal scene graph generation is an essential task in household activity recognition that aims to identify human-object interactions. Constructing a dataset with per-frame object region and consistent relationship annotations requires extremely high labor costs. Existing datasets sparsely annotate frames sampled from videos, resulting in the lack of dense spatio-temporal correlation in videos. Additionally, existing datasets contain inconsistent relationship annotations, leading to the problem of learning ambiguous temporal associations. Moreover, existing datasets mainly discuss relationships that can be inferred from a single frame, ignoring the significance of temporal associations. To resolve those issues, we created a simulated dataset with per-frame consistent annotations and introduced a range of relationships requiring both spatial and temporal context. Most existing methods explore spatial correlations within single images and do not explicitly consider the dynamic changes across frames. Therefore, we proposed a tracking-based approach that explicitly grasps spatio-temporal human-object interactions while simultaneously localizing humans and objects. Our proposed approach achieved state-of-the-art performance on scene graph generation and outperformed existing methods in scene graph localization by large margins on the proposed dataset. Moreover, the experiments show the efficacy of pre-training on the proposed dataset while adapting to a previous benchmark consisting of real daily videos, indicating the potential of the proposed dataset in real-world scenarios. Yue Qiu 0001, Yoshiki Nagasaki, Kensho Hara, Hirokatsu Kataoka, Ryota Suzuki 0006, Kenji Iwata, Yutaka Satoh |
WACV | 4 |
| 2023 | 3D Change Localization and Captioning from Dynamic Scans of Indoor ScenesabstractDaily indoor scenes often involve constant changes due to human activities. To recognize scene changes, existing change captioning methods focus on describing changes from two images of a scene. However, to accurately perceive and appropriately evaluate physical changes and then identify the geometry of changed objects, recognizing and localizing changes in 3D space is crucial. Therefore, we propose a task to explicitly localize changes in 3D bounding boxes from two point clouds and describe detailed scene changes, including change types, object attributes, and spatial locations. Moreover, we create a simulated dataset with various scenes, allowing generating data without labor costs. We further propose a framework that allows different 3D object detectors to be incorporated in the change detection process, after which captions are generated based on the correlations of different change regions. The proposed framework achieves promising results in both change detection and captioning. Furthermore, we also evaluated on data collected from real scenes. The experiments show that pretraining on the proposed dataset increases the change detection accuracy by +12.8% (mAP0.25) when applied to real-world data. We believe that our proposed dataset and discussion could provide both a new benchmark and in-sights for future studies in scene change understanding. Yue Qiu 0001, Shintaro Yamamoto, Ryosuke Yamada, Ryota Suzuki 0006, Hirokatsu Kataoka, Kenji Iwata, Yutaka Satoh |
WACV | 5 |
| 2022 | Can Vision Transformers Learn without Natural Images?abstractIs it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated labels, the recent use of natural images has resulted in problems related to privacy violation, inadequate fairness protection, and the need for labor-intensive annotations. In this paper, we experimentally verify that the results of formula-driven supervised learning (FDSL) framework are comparable with, and can even partially outperform, sophisticated self-supervised learning (SSL) methods like SimCLRv2 and MoCov2 without using any natural images in the pre-training phase. We also consider ways to reorganize FractalDB generation based on our tentative conclusion that there is room for configuration improvements in the iterated function system (IFS) parameter settings of such databases. Moreover, we show that while ViTs pre-trained without natural images produce visualizations that are somewhat different from ImageNet pre-trained ViTs, they can still interpret natural image datasets to a large extent. Finally, in experiments using the CIFAR-10 dataset, we show that our model achieved a performance rate of 97.8, which is comparable to the rate of 97.4 achieved with SimCLRv2 and 98.0 achieved with ImageNet. Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata, Nakamasa Inoue, Yutaka Satoh |
AAAI | 2 |
| 2022 | Replacing Labeled Real-image Datasets with Auto-generated ContoursabstractIn the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k without the use of real images, human-, and self-supervision during the pre-training of Vision Transformers (ViTs). For example, ViT-Base pre-trained on ImageNet-21k shows 81.8% top-1 accuracy when fine-tuned on ImageNet-1k and FDSL shows 82.7% top-1 accuracy when pre-trained under the same conditions (number of images, hyperparameters, and number of epochs). Images generated by formulas avoid the privacy/copyright issues, labeling cost and errors, and biases that real images suffer from, and thus have tremendous potential for pre-training general models. To understand the performance of the synthetic images, we tested two hypotheses, namely (i) object contours are what matter in FDSL datasets and (ii) increased number of parameters to create labels affects performance improvement in FDSL pre-training. To test the former hypothesis, we constructed a dataset that consisted of simple object contour combinations. We found that this dataset can match the performance of fractals. For the latter hypothesis, we found that increasing the difficulty of the pre-training task generally leads to better fine-tuning accuracy. Hirokatsu Kataoka, Ryo Hayamizu, Ryosuke Yamada, Kodai Nakashima, Sora Takashima, Edgar Josafat Martinez-Noriega, Nakamasa Inoue, Rio Yokota |
CVPR | 1 |
| 2022 | Point Cloud Pre-training with Natural 3D StructuresabstractThe construction of 3D point cloud datasets requires a great deal of human effort. Therefore, constructing a large-scale 3D point clouds dataset is difficult. In order to rem-edy this issue, we propose a newly developed point cloud fractal database (PC-FractalDB), which is a novel family of formula-driven supervised learning inspired by fractal geometry encountered in natural 3D structures. Our re-search is based on the hypothesis that we could learn rep-resentations from more real-world 3D patterns than con-ventional 3D datasets by learning fractal geometry. We show how the PC-FractalDB facilitates solving several re-cent dataset-related problems in 3D scene understanding, such as 3D model collection and labor-intensive annotation. The experimental section shows how we achieved the performance rate of up to 61.9% and 59.0% for the Scan-NetV2 and SUN RGB-D datasets, respectively, over the current highest scores obtained with the PointContrast, con-trastive scene contexts (CSC), and RandomRooms. More-over, the PC-FractalDB pre-trained model is especially ef-fective in training with limited data. For example, in 10% of training data on ScanNetV2, the PC-FractalDB pre-trained VoteNet performs at 38.3%, which is +14.8% higher accu-racy than CSC. Of particular note, we found that the pro-posed method achieves the highest results for 3D object de-tection pre-training in limited point cloud data.11Dataset release: https://ryosuke-yamada.github.io/PointCloud-FractalDataBase/ Ryosuke Yamada, Hirokatsu Kataoka, Naoya Chiba, Yukiyasu Domae, Tetsuya Ogata |
CVPR | 2 |
| 2022 | Neural Density-Distance Fields
Itsuki Ueda, Yoshihiro Fukuhara, Hirokatsu Kataoka, Hiroaki Aizawa, Hidehiko Shishido, Itaru Kitahara |
ECCV (32) | 3 |
| 2022 | Spatiotemporal Initialization for 3D CNNs with Generated Motion PatternsabstractThe paper proposes a framework of Formula-Driven Supervised Learning (FDSL) for spatiotemporal initialization. Our FDSL approach enables to automatically and simultaneously generate motion patterns and their video labels with a simple formula which is based on Perlin noise. We designed a dataset of generated motion patterns adequate for the 3D CNNs to learn a better basis set of natural videos. The constructed Video Perlin Noise (VPN) dataset can be applied to initialize a model before pre-training with large-scale video datasets such as Kinetics-400/700, to enhance target task performance. Our spatiotemporal initialization with VPN dataset (VPN initialization) outperforms the previous initialization method with the inflated 3D ConvNet (I3D) using 2D ImageNet dataset. Our proposed method increased the top-1 video-level accuracy of Kinetics-400 pre-trained model on {Kinetics-400, UCF-101, HMDB-51, ActivityNet} datasets. Especially, the proposed method increased the performance rate of Kinetics-400 pre-trained model by 10.3 pt on ActivityNet. We also report that the relative performance improvements from the baseline are greater in 3D CNNs rather than other models. Our VPN initialization mainly helps to enhance the performance in spatiotemporal 3D kernels. The datasets, codes and pre-trained models used in this study will be publicly available1. Hirokatsu Kataoka, Kensho Hara, Ryusuke Hayashi, Eisuke Yamagata, Nakamasa Inoue |
WACV | 1 |
| 2022 | Pre-Training Without Natural ImagesabstractAbstract Is it possible to use convolutional neural networks pre-trained without any natural images to assist natural image understanding? The paper proposes a novel concept, Formula-driven Supervised Learning (FDSL). We automatically generate image patterns and their category labels by assigning fractals, which are based on a natural law. Theoretically, the use of automatically generated images instead of natural images in the pre-training phase allows us to generate an infinitely large dataset of labeled images. The proposed framework is similar yet different from Self-Supervised Learning because the FDSL framework enables the creation of image patterns based on any mathematical formulas in addition to self-generated labels. Further, unlike pre-training with a synthetic image dataset, a dataset under the framework of FDSL is not required to define object categories, surface texture, lighting conditions, and camera viewpoint. In the experimental section, we find a better dataset configuration through an exploratory study, e.g., increase of #category/#instance, patch rendering, image coloring, and training epoch. Although models pre-trained with the proposed Fractal DataBase (FractalDB), a database without natural images, do not necessarily outperform models pre-trained with human annotated datasets in all settings, we are able to partially surpass the accuracy of ImageNet/Places pre-trained models. The FractalDB pre-trained CNN also outperforms other pre-trained models on auto-generated datasets based on FDSL such as Bezier curves and Perlin noise. This is reasonable since natural objects and scenes existing around us are constructed according to fractal geometry. Image representation with the proposed FractalDB captures a unique feature in the visualization of convolutional layers and attentions. Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, Yutaka Satoh |
Int. J. Comput. Vis. | 1 |
| 2022 | Predicting Appearance of Vehicles From Blind Spots Based on Pedestrian Behaviors at CrossroadsabstractConventional prediction approaches for traffic scenes primarily predict the future states of visible objects (i.e., not in blind spots) based on their current observations. This study focused on the prediction of future states of objects in blind spots (e.g., those outside the filed-of-view or occluded regions) based on the current observations of other visible objects. We proposed a method that predicts the appearance of vehicles from a blind spot based on the behaviors of visible pedestrians who observe vehicles in the blind spot. Our proposed method utilizes a spatiotemporal 3D convolutional neural network and learns pedestrian behaviors for predictions. The method explicitly represents subtle motions and the surrounding environments of pedestrians using pose estimation and semantic segmentation. To conduct evaluation experiments, we built two datasets of videos capturing real traffic scenes. The datasets are collected by cameras with and without ego-motions. Using the datasets, we conducted experiments not only on simpler configurations but also on realistic traffic environments. Based on the experimental results, the following conclusions could be obtained: (i) our proposed method achieved a high performance at a level similar to that of humans in our prediction task, and predicted the appearance of vehicles from blind spots more than 1.5 s before they actually appeared. (ii) Explicit representations of pose and semantic masks captured information complementary to RGB videos, and ensembling the representations improved the prediction performance. (iii) Fine-tuning the models using videos with ego-motions is important to achieve good prediction in the videos captured by driving cars. Kensho Hara, Hirokatsu Kataoka, Masaki Inaba, Kenichi Narioka, Ryusuke Hotta, Yutaka Satoh |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Describing and Localizing Multiple Changes with TransformersabstractChange captioning tasks aim to detect changes in image pairs observed before and after a scene change and generate a natural language description of the changes. Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-change captioning dataset; (ii) We benchmark existing state-of-the-art methods of single change captioning on multi-change captioning; (iii) We further propose Multi-Change Captioning transformers (MCCFormers) that identify change regions by densely correlating different regions in image pairs and dynamically determines the related change regions with words in sentences. The proposed method obtained the highest scores on four conventional change captioning evaluation metrics for multi-change captioning. Additionally, our proposed method can separate attention maps for each change and performs well with respect to change localization. Moreover, the proposed framework outperformed the previous state-of-the-art methods on an existing change captioning benchmark, CLEVR-Change, by a large margin (+6.1 on BLEU-4 and +9.7 on CIDEr scores), indicating its general ability in change captioning tasks. The code and dataset are available at the project page1. Yue Qiu 0001, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki 0006, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh |
ICCV | 6 |
| 2021 | MV-FractalDB: Formula-driven Supervised Learning for Multi-view Image RecognitionabstractThe paper proposes a method for automatic multi-view dataset construction based on formula-driven supervised learning (FDSL). Although data collection and human annotation of 3D objects are labor-intensive, we automatically generate their training data and labels in the proposed multi-view dataset. To create a large-scale multi-view dataset, we employ fractal geometry, which is considered the background information of many objects in the real world. We project in a circle from the rendered 3D fractal models to construct the Multi-view Fractal DataBase (MV-FractalDB), which is then used to make a pre-trained CNN model. According to the experimental results, the MV-FractalDB pre-trained model surpasses the accuracies with self-supervised methods (e.g., SimCLR and MoCo) and is close to supervised methods (e.g., ImageNet) in terms of performance rates on multi-view image datasets. We demonstrate the potential of FDSL for multi-view image recognition. Ryosuke Yamada, Ryota Suzuki 0006, Akio Nakamura, Yusuke Yoshiyasu, Ryusuke Sagawa, Hirokatsu Kataoka |
IROS | 7 |
| 2021 | Viewpoint-agnostic Image RenderingabstractRendering an any-viewpoint image is extremely difficult for Generative Adversarial Networks. This is because conventional GANs do not understand 3D information under-lying a given viewpoint image such as an object shape and relationship between viewpoint and objects in 3D space. In this paper, we present how to perform a Viewpoint-Agnostic Image Rendering (VAIR), equipping a conditional GAN with a mechanism to reconstruct 3D information of the input view. VAIR realizes any-viewpoint image generation by manipulating a viewpoint in 3D space where the reconstructed instance shape is arranged. In addition, we convert the reconstructed 3D shape into a 2D representation for image-based conditional GAN, while preserving detail 3D information. The representation consists of a depth image and 2D semantic keypoint images, which are obtained by rendering the shape from a viewpoint. In the experiment, we evaluate using a CUB-200-2011 dataset, which contains few-samples biased a viewpoint such that covers only part of the target appearance. As a result, our VAIR clearly renders an any-viewpoint image. Hiroaki Aizawa, Hirokatsu Kataoka, Yutaka Satoh, Kunihito Kato |
WACV | 2 |
| 2021 | Alleviating Over-segmentation Errors by Detecting Action BoundariesabstractWe propose an effective framework for the temporal action segmentation task, namely an Action Segment Refinement Framework (ASRF). Our model architecture consists of a long-term feature extractor and two branches: the Action Segmentation Branch (ASB) and the Boundary Regression Branch (BRB). The long-term feature extractor provides shared features for the two branches with a wide temporal receptive field. The ASB classifies video frames with action classes, while the BRB regresses the action boundary probabilities. The action boundaries predicted by the BRB refine the output from the ASB, which results in a significant performance improvement. Our contributions are three-fold: (i) We propose a framework for temporal action segmentation, the ASRF, which divides temporal action segmentation into frame-wise action classification and action boundary regression. Our framework refines frame-level hypotheses of action classes using predicted action boundaries. (ii) We propose a loss function for smoothing the transition of action probabilities, and analyze combinations of various loss functions for temporal action segmentation. (iii) Our framework outperforms state-of-the-art methods on three challenging datasets, offering an improvement of up to 13.7% in terms of segmental edit distance and up to 16.1% in terms of segmental F1 score. Our code is publicly available1. Yuchi Ishikawa, Seito Kasai, Yoshimitsu Aoki, Hirokatsu Kataoka |
WACV | 4 |
| 2020 | Pre-training Without Natural Images
Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, Yutaka Satoh |
ACCV (6) | 1 |
| 2020 | Retrieving and Highlighting Action with Spatiotemporal ReferenceabstractIn this paper, we present a framework that jointly retrieves and spatiotemporally highlights actions in videos by enhancing current deep cross-modal retrieval methods. Our work takes on the novel task of action highlighting, which visualizes where and when actions occur in an untrimmed video setting. Action highlighting is a fine-grained task, compared to conventional action recognition tasks which focus on classification or window-based localization. Leveraging weak supervision from annotated captions, our framework acquires spatiotemporal relevance maps and generates local embeddings which relate to the nouns and verbs in captions. Through experiments, we show that our model generates various maps conditioned on different actions, in which conventional visual reasoning methods only go as far as to show a single deterministic saliency map. Also, our model improves retrieval recall over our baseline without alignment by 2-3% on the MSR-VTT dataset. Seito Kasai, Yuchi Ishikawa, Masaki Hayashi, Yoshimitsu Aoki, Kensho Hara, Hirokatsu Kataoka |
ICIP | 6 |
| 2020 | Disentangle, Assemble, and Synthesize: Unsupervised Learning to Disentangle Appearance and LocationabstractThe next step for the generative adversarial networks (GAN) is to learn representations that allow us to control only a certain factor in the image explicitly. Since such a representation of the factor is independent of other factors, the controllability obtained from these representations leads to interpretability by identifying the variation of the synthesized image and the transferability for downstream tasks by inference. However, since it is difficult to identify and strictly define latent factors, the annotation is laborious. Moreover, learning such representations by a GAN is challenging due to the complex generation process. Therefore, we resolve this limitation using a novel generative model that can disentangle latent space into the appearance, the x-axis, and the y-axis of the object, and reassemble these components in an unsupervised manner. Specifically, based on the concept of packing the appearance and location in each position of the feature map, we introduce a novel structural constraint technique that prevents these representations from interacting with each other. The proposed structural constraint promotes the disentanglement of these factors. In experiments, we found that the proposed method is simple but effective for controllability and allows us to control the appearance and location via latent space without supervision, as compared with the conditional GAN. Hiroaki Aizawa, Hirokatsu Kataoka, Yutaka Satoh, Kunihito Kato |
ICPR | 2 |
| 2020 | Initialization Using Perlin Noise for Training Networks with a Limited Amount of DataabstractWe propose a novel network initialization method using Perlin noise for training image classification networks with a limited amount of data. Our main idea is to initialize the network parameters by solving an artificial noise classification problem, where the aim is to classify Perlin noise samples into their noise categories. Specifically, the proposed method consists of two steps. First, it generates Perlin noise samples with category labels defined based on noise complexity. Second, it solves a classification problem, in which network parameters are optimized to classify the generated noise samples. This method produces a reasonable set of initial weights (filters) for image classification. To the best of our knowledge, this is the first work to initialize networks by solving an artificial optimization problem without using any real-world images. Our experiments show that the proposed method outperforms conventional initialization methods on four image classification datasets. Nakamasa Inoue, Eisuke Yamagata, Hirokatsu Kataoka |
ICPR | 3 |
| 2020 | Augmented Cyclic Consistency Regularization for Unpaired Image-to-Image TranslationabstractUnpaired image-to-image (I2I) translation has received considerable attention in pattern recognition and computer vision because of recent advancements in generative adversarial networks (GANs). However, due to the lack of explicit supervision, unpaired I2I models often fail to generate realistic images, especially in challenging datasets with different backgrounds and poses. Hence, stabilization is indispensable for GANs and applications of I2I translation. Herein, we propose Augmented Cyclic Consistency Regularization (ACCR), a novel regularization method for unpaired I2I translation. Our main idea is to enforce consistency regularization originating from semi-supervised learning on the discriminators leveraging real, fake, reconstructed, and augmented samples. We regularize the discriminators to output similar predictions when fed pairs of original and perturbed images. We qualitatively clarify why consistency regularization on fake and reconstructed samples works well. Quantitatively, our method outperforms the consistency regularized GAN (CR-GAN) in real-world translations and demonstrates efficacy against several data augmentation variants and cycle-consistent constraints. Takehiko Ohkawa, Naoto Inoue, Hirokatsu Kataoka, Nakamasa Inoue |
ICPR | 3 |
| 2020 | Adversarial Knowledge Distillation for a Compact GeneratorabstractIn this paper, we propose memory-efficient Generative Adversarial Nets (GANs) in line with knowledge distillation. Most existing GANs have a shortcoming in terms of the number of model parameters and low processing speed. Here, to tackle the problem, we propose Adversarial Knowledge Distillation for Generative models (AKDG) for highly efficient GANs, in terms of unconditional generation. Using AKDG, model size and processing speed are substantively reduced. Through an adversarial training exercise with a distillation discriminator, a student generator successfully mimics a teacher generator in fewer model layers and fewer parameters and at a higher processing speed. Moreover, our AKDG is network architecture-agnostic. A Comparison of AKDG-applied models to vanilla models suggests that it achieves closer scores to a teacher generator and more efficient performance than a baseline method with respect to Inception Score (IS) and Frechet Inception Distance (FID). In CIFAR-10 experiments, improving IS/FID 1.17pt/55.19pt and in LSUN bedroom experiments, improving FID 71.1pt in comparison to the conventional distillation method for GANs. Our project page is https://maguro27.github.io/AKDG/. Hideki Tsunashima, Hirokatsu Kataoka, Junji Yamato, Qiu Chen, Shigeo Morishima |
ICPR | 2 |
| 2020 | Joint Pedestrian Detection and Risk-level Prediction with Motion-Representation-by-DetectionabstractThe paper presents a pedestrian near-miss detector with temporal analysis that provides both pedestrian detection and risk-level predictions which are demonstrated on a self-collected database. Our work makes three primary contributions: (i) The framework of pedestrian near-miss detection is proposed by providing both a pedestrian detection and risk-level assignment. Specifically, we have created a Pedestrian Near-Miss (PNM) dataset that categorizes traffic near-miss incidents based on their risk levels (high-, low-, and no-risk). Unlike existing databases, our dataset also includes manually localized pedestrian labels as well as a large number of incident-related videos. (ii) Single-Shot MultiBox Detector with Motion Representation (SSD-MR) is implemented to effectively extract motion-based features in a detected pedestrian. (iii) Using the self-collected PNM dataset and SSD-MR, our proposed method achieved +19.38% (on risk-level prediction) and +13.00% (on joint pedestrian detection and risk-level prediction) higher scores than that of the baseline SSD and LSTM. Additionally, the running time of our system is over 50 fps on a graphics processing unit (GPU). Hirokatsu Kataoka, Teppei Suzuki, Kodai Nakashima, Yutaka Satoh, Yoshimitsu Aoki |
ICRA | 1 |
| 2019 | Incorporating 3D Information Into Visual Question AnsweringabstractWe propose a tactic of advancing Visual Question Answering (VQA) task by incorporating 3D information via multi-view images. Conventional VQA approaches, which reply an answer in words against a linguistic question about a given RGB image, have less ability to recognize geometrical information so that they tend to fail to count things or guess positional relationship. Moreover, they have no ability to determine blinded space, so it is not feasible to invent VQA function to robots which will work in highly-occluded real-world environments. To achieve the situation, we introduce a new multi-view VQA dataset along with an approach that incorporating 3D scene information directly captured from multi-view images into VQA without using depth images or employing SLAM. Our proposed approach achieves strong performance with an overall accuracy of 95.4% on the challenging multi-view VQA dataset setup, which contains relatively severe occlusion. This work also demonstrates the promising aspects of bridging the gap between 3D vision and language. Yue Qiu 0001, Yutaka Satoh, Ryota Suzuki 0006, Hirokatsu Kataoka |
3DV | 4 |
| 2019 | Unsupervised Out-of-context Action UnderstandingabstractThe paper presents an unsupervised out-of-context action (O2CA) paradigm that is based on facilitating understanding by separately presenting both human action and context within a video sequence. As a means of generating an unsupervised label, we comprehensively evaluate responses from action-based (ActionNet) and context-based (ContextNet) convolutional neural networks (CNNs). Additionally, we have created three synthetic databases based on the human action (UCF101, HMDB51) and motion capture (mocap) (SURREAL) datasets. We then conducted experimental comparisons between our approach and conventional approaches. We also compared our unsupervised learning method with supervised learning using an O2CA ground truth given by synthetic data. From the results obtained, we achieved a 96.8 score on Synth-UCF, a 96.8 score on Synth-HMDB, and 89.0 on SURREAL-O2CA with F-score. Hirokatsu Kataoka, Yutaka Satoh |
ICRA | 1 |
| 2018 | Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?abstractThe purpose of this study is to determine whether current video datasets have sufficient data for training very deep convolutional neural networks (CNNs) with spatio-temporal three-dimensional (3D) kernels. Recently, the performance levels of 3D CNNs in the field of action recognition have improved significantly. However, to date, conventional research has only explored relatively shallow 3D architectures. We examine the architectures of various 3D CNNs from relatively shallow to very deep ones on current video datasets. Based on the results of those experiments, the following conclusions could be obtained: (i) ResNet-18 training resulted in significant overfitting for UCF-101, HMDB-51, and ActivityNet but not for Kinetics. (ii) The Kinetics dataset has sufficient data for training of deep 3D CNNs, and enables training of up to 152 ResNets layers, interestingly similar to 2D ResNets on ImageNet. ResNeXt-101 achieved 78.4% average accuracy on the Kinetics test set. (iii) Kinetics pretrained simple 3D architectures outperforms complex 2D architectures, and the pretrained ResNeXt-101 achieved 94.5% and 70.2% on UCF-101 and HMDB-51, respectively. The use of 2D CNNs trained on ImageNet has produced significant progress in various tasks in image. We believe that using deep 3D CNNs together with Kinetics will retrace the successful history of 2D CNNs and ImageNet, and stimulate advances in computer vision for videos. The codes and pretrained models used in this study are publicly available1. Kensho Hara, Hirokatsu Kataoka, Yutaka Satoh |
CVPR | 2 |
| 2018 | Anticipating Traffic Accidents With Adaptive Loss and Large-Scale Incident DBabstractIn this paper, we propose a novel approach for traffic accident anticipation through (i) Adaptive Loss for Early Anticipation (AdaLEA) and (ii) a large-scale self-annotated incident database for anticipation. The proposed AdaLEA allows a model to gradually learn an earlier anticipation as training progresses. The loss function adaptively assigns penalty weights depending on how early the model can anticipate a traffic accident at each epoch. Additionally, we construct a Near-miss Incident DataBase for anticipation. This database contains an enormous number of traffic near-miss incident videos and annotations for detail evaluation of two tasks, risk anticipation and risk-factor anticipation. In our experimental results, we found our proposal achieved the highest scores for risk anticipation (+6.6% better on mean average precision (mAP) and 2.36 sec earlier than previous work on the average time-to-collision (ATTC)) and risk-factor anticipation (+4.3% better on mAP and 0.70 sec earlier than previous work on ATTC). Hirokatsu Kataoka, Yoshimitsu Aoki, Yutaka Satoh |
CVPR | 2 |
| 2018 | Fashion Culture Database: Construction of Database for World-wide Fashion AnalysisabstractThe paper presents a novel concept that analyzes and visualizes worldwide fashion styles. Our goal is to reveal web-based viral fashion styles. To achieve the fashion-based analysis, we have collected fashion culture database (FCDB), which consists of 76 million geo-tagged images in 16 cosmopolitan cities. The database allows us to grasp a trend of mixed fashion styles with a fashion-based descriptor and codeword vector. In order to unveil web-based fashion trends in the FCDB, we applied a simple technique that is a temporal subtraction between consecutive codeword vectors in two different times. In the experiments, we show the analysis of fashion trends and fashion-based city similarity in a social media. As the result of large-scale data collection, we achieved world-level fashion visualization. Kaori Abe, Munetaka Minoguchi, Teppei Suzuki, Naofumi Akimoto, Yue Qiu 0001, Ryota Suzuki 0006, Kenji Iwata, Yutaka Satoh, Hirokatsu Kataoka |
ICARCV | 10 |
| 2018 | Semantic Change DetectionabstractChange detection is the study of detecting changes between two different images of a scene taken at different times. The change detection methodology can provide us information in which area images changed time by time. However, for application use, especially on disaster investigation, it is highly required to understand not only where but also what changes are occurred in high precision and resolution. The paper proposes the concept of semantic change detection, which involves intuitively inserting semantic meaning into detected change areas. We mainly focus on the novel semantic segmentation in addition to a conventional change detection approach. In order to solve this problem and obtain a high-level of performance, we propose an improvement to the hypercolumns representation, hereafter known as hypermaps, which effectively uses convolutional maps obtained from convolutional neural networks (CNNs). We also employ multi-scale feature representation captured by different image patches. We applied our method to the TSUNAMI panoramic change detection dataset (TSUNAMI dataset), and re-annotated the changed areas of the dataset via semantic classes. The results show that our multi-scale hypermaps provided outstanding performance on the re-annotated TSUNAMI dataset. Munetaka Minoguchi, Ryota Suzuki 0006, Akio Nakamura, Kenji Iwata, Yutaka Satoh, Hirokatsu Kataoka |
ICARCV | 7 |
| 2018 | Automatic paper summary generation from visual and textual informationabstractDue to the recent boom in artificial intelligence (AI) research, including computer vision (CV), it has become impossible for researchers in these fields to keep up with the exponentially increasing number of manuscripts. In response to this situation, this paper proposes the paper summary generation (PSG) task using a simple but effective method to automatically generate an academic paper summary from raw PDF data. We realized PSG by combination of vision-based supervised components detector and language-based unsupervised important sentence extractor, which is applicable for a trained format of manuscripts. We show the quantitative evaluation of ability of simple vision-based components extraction, and the qualitative evaluation that our system can extract both visual item and sentence that are helpful for understanding. After processing via our PSG, the 979 manuscripts accepted by the Conference on Computer Vision and Pattern Recognition (CVPR) 2018 are available1 . It is believed that the proposed method will provide a better way for researchers to stay caught with important academic papers. Shintaro Yamamoto, Yoshihiro Fukuhara, Ryota Suzuki 0006, Shigeo Morishima, Hirokatsu Kataoka |
ICMV | 5 |
| 2018 | Towards Good Practice for Action Recognition with Spatiotemporal 3D ConvolutionsabstractThe purpose of this study is to explore good practice for training convolutional neural networks (CNNs) with spatiotemporal three-dimensional (3D) kernels. Recently, 3D CNNs in the field of action recognition are rapidly developed, and the performance levels of them have improved significantly. However, to date, conventional research has mainly focused on their architecture, and has not sufficiently explored their training configurations. We conduct various experiments with different training configurations on Kinetics, UCF-101, and HMDB-51 datasets to share the knowledge of 3D CNNs for the research community. According to the results of those experiments, the following conclusions could be obtained. (i) Data augmentation by spatiotemporal random cropping improved the performance levels. (ii) Data augmentation by multi-scale spatial cropping increased the accuracies in most cases whereas multi-scale temporal cropping decreased them. (iii) A corner cropping strategy, which is previously shown as a good method for two-stream 2D CNNs, resulted lower accuracies for 3D CNNs compared with simple random cropping. (iv) Freezing early layers of 3D CNNs improved the performance levels when fine-tuning 3D CNNs on a relatively small dataset. Kensho Hara, Hirokatsu Kataoka, Yutaka Satoh |
ICPR | 2 |
| 2018 | Occlusion Handling Human Detection with Refocused ImagesabstractThe paper presents a novel robust human detection method based on camera array system to broaden the application range for human detection. Currently, even by using a deep neural network (DNN), it is difficult to detect a hardly occluded human. In the camera array system, we consider how to distinctly show a human occluded by an environmental condition. The generated refocused images by the camera array system allow us to remove the effect of the noises. Although refocused images have not been utilized in conventional human detection, we believe that the refocused images are beneficial for improving the detection performance, especially in severe conditions. To execute the experiments, we have collected Refocused Human DataBase (RHDB) with the camera array system. By using HOG+SVM with a monocular camera (at an almost random rate of 54.8%), the refocused images made the +10.1% improvement (64.9%) by noticeably showing a human. The combined representation of refocused images and AlexNet achieved 94.6% on the RHDB. Moreover, our final model recorded 98.0% with an attention-layer and fine-tuned parameters. Hirokatsu Kataoka, Shuhei Ohki, Kenji Iwata, Yutaka Satoh |
ICPR | 1 |
| 2018 | Drive Video Analysis for the Detection of Traffic Near-Miss IncidentsabstractBecause of their recent introduction, self-driving cars and advanced driver assistance system (ADAS) equipped vehicles have had little opportunity to learn, the dangerous traffic (including near-miss incident) scenarios that provide normal drivers with strong motivation to drive safely. Accordingly, as a means of providing learning depth, this paper presents a novel traffic database that contains information on a large number of traffic near-miss incidents that were obtained by mounting driving recorders in more than 100 taxis over the course of a decade. The study makes the following two main contributions: (i) In order to assist automated systems in detecting near-miss incidents based on database instances, we created a large-scale traffic near-miss incident database (NIDB) that consists of video clip of dangerous events captured by monocular driving recorders. (ii) To illustrate the applicability of NIDB traffic near-miss incidents, we provide two primary database-related improvements: parameter fine-tuning using various near-miss scenes from NIDB, and foreground/background separation into motion representation. Then, using our new database in conjunction with a monocular driving recorder, we developed a near-miss recognition method that provides automated systems with a performance level that is comparable to a human-level understanding of near-miss incidents (64.5% vs. 68.4% at near-miss recognition, 61.3% vs. 78.7% at near-miss detection). Hirokatsu Kataoka, Teppei Suzuki, Shoko Oikawa, Yasuhiro Matsui, Yutaka Satoh |
ICRA | 1 |
| 2017 | Text Detection in Traffic Informatory Signs Using Synthetic DataabstractTraffic informatory signs, which is a category of traffic signs and text-based signs, is very important to both drivers and intelligent transport systems. Previous studies have usually sought to extract text lines in signs to apply to optical character recognition (OCR) system, but they do not work well in real-world conditions with severe disturbance. In this paper, we report on our study of place name text detection and recognition on traffic informatory signs using convolutional neural networks (CNNs) and transform traditional text detection and recognition into word-level multi-class robust image classification. In our study, each place name corresponds to one class. Because the number of word classes is large and collecting real images for training dataset is difficult, we generate several synthetic datasets mainly by means of two methods and use them to train the CNN respectively. One method generates with standard templates of the signs, while the other method renders text in natural images which have no relationship with the signs. Our experimental results show that our method can achieve high levels of accuracy when reading traffic informatory signs in real-world conditions. Accuracy of 0.891 and 0.981 are achieved in the former and latter method of generating dataset, which verify that proposed methods are effective in our task. We also analyze the dependence of color channels in text detection during our task, which help generate the more efficacious synthetic dataset. Fangge Chen, Hirokatsu Kataoka, Yutaka Satoh |
ICDAR | 2 |
| 2016 | Recognition of Transitional Action for Short-Term Action Prediction using Discriminative Temporal CNN Feature
Hirokatsu Kataoka, Yudai Miyashita, Masaki Hayashi, Kenji Iwata, Yutaka Satoh |
BMVC | 1 |
| 2015 | Predicting driving behavior using inverse reinforcement learning with multiple reward functions towards environmental diversityabstractPredicting defensive driving is a promising technology for novel advanced driver assistance systems. In recent years, modeling driving behavior in residential roads through inverse reinforcement learning (IRL) has been attracting attention in intelligent vehicle community thanks to the superiority of this approach providing long-term prediction of fine-grained driving behavior. However, it suffers from poor performance in diverse environment due to the fact that the single reward function could not handle all the environment with large diversity. Towards this issue, a novel IRL framework with multiple reward functions to deal with environmental diversity is proposed in the paper. Specifically, the model employs Dirichlet process mixtures as a flexible and powerful Bayesian model to divide the environment into clusters and learns the parameters in each cluster simultaneously. Experimental result with expert driver behavior data shows that our model with multiple reward functions provides superior performance over the IRL model with single reward function. It also suggests that the clustering of environments based on the driving behavior of professional drivers could be useful on evaluating driving environments. Masamichi Shimosaka, Kentaro Nishi, Jun-ichi Satoh, Hirokatsu Kataoka |
Intelligent Vehicles Symposium | 4 |
| 2015 | What is an Effective Feature for a Detection Problem? Feature Evaluation in Multiple ScenesabstractWe investigated effective features for human detection. The histogram of oriented gradients (HOG), which was proposed by N. Dalal, is an important representation that accumulates the edge-magnitude into a quantized histogram. Effective features similar to the HOG have been proposed. We question what the most effective feature is. We thus evaluate several features on three datasets of pedestrians, faces, and vehicles. We select the scale-invariant feature transform, local binary pattern, higher-order local auto correlation (HLAC), co-occurrence HOG, and extended CoHOG in addition to the HOG as features. These features have been adopted as effective features in related works. The features are applied to human detection on each dataset employing the real AdaBoost classifier. A comparison of classification results reveals that the combination of the HLAC and CoHOG is an effective feature for human detection. Tomoaki K. Yamabe, Yudai Miyashita, Shin-ichi Sato, Yudai Yamamoto, Akio Nakamura, Hirokatsu Kataoka |
SMC | 6 |
| 2014 | Extended Co-occurrence HOG with Dense Trajectories for Fine-Grained Activity Recognition
Hirokatsu Kataoka, Kiyoshi Hashimoto, Kenji Iwata, Yutaka Satoh, Nassir Navab, Slobodan Ilic, Yoshimitsu Aoki |
ACCV (5) | 1 |
| 2014 | Feature integration with random forests for real-time human activity recognitionabstractThis paper presents an approach for real-time human activity recognition. Three different kinds of features (flow, shape, and a keypoint-based feature) are applied in activity recognition. We use random forests for feature integration and activity classification. A forest is created at each feature that performs as a weak classifier. The international classification of functioning, disability and health (ICF) proposed by WHO is applied in order to set the novel definition in activity recognition. Experiments on human activity recognition using the proposed framework show - 99.2% (Weizmann action dataset), 95.5% (KTH human actions dataset), and 54.6% (UCF50 dataset) recognition accuracy with a real-time processing speed. The feature integration and activity-class definition allow us to accomplish high-accuracy recognition match for the state-of-the-art in real-time. Hirokatsu Kataoka, Kiyoshi Hashimoto, Yoshimitsu Aoki |
ICMV | 1 |
| 2013 | Robust human tracking using statistical human shape model with postural variationabstractHuman tracking in monocular image sequences has been studied in the field of computer vision for many kinds of applications such as surveillance system, intelligent room, sports video analysis and so on. Human tracking in real environment is challenging topic due to various factors such as illumination change, partial or almost complete occlusion of human body, and wide variety of body shapes. In this paper, we present a robust human tracking using statistical human shape model of appearance variation with postural change. Our part-based statistical human model can generate learned appearances of main human poses, and enables effective and robust human tracking with simple features such silhouette, edge and color. Our proposed method achieves human tracking robust not only to partial occlusion but also to postural change. The experimental results validate the robustness of our methods in the real indoor environments. Kiyoshi Hashimoto, Hirokatsu Kataoka, Yoshimitsu Aoki, Yuji Sato |
IECON | 2 |
| 2013 | Robust feature descriptor and vehicle motion model with tracking-by-detection for active safetyabstractThe percentage of pedestrian deaths in traffic accidents is on the rise in Japan. In recent years, there have been calls for measures to be introduced to protect vulnerable road users such as pedestrians and cyclists. In this study, a method to detect and track pedestrians using an in-vehicle camera is presented to perform braking controls, warn the driver, and develop improved safety systems for pedestrians. We improved the technology of detecting pedestrians using highly accurate images obtained with a monocular camera. We were able to predict pedestrian activity by monitoring the images, and developed an algorithm with which to recognize pedestrians and their movements more accurately. The effectiveness of the algorithm was tested using images taken on real roads. For the feature descriptor, we used an extended co-occurrence histogram of oriented gradients (ECoHOG) that accumulated the integration of gradient intensities. In the tracking step, we applied an effective motion model using optical flow and the proposed feature descriptor ECoHOG in a tracking-by-detection framework. These techniques were verified using images captured on the real road. Hirokatsu Kataoka, Kimimasa Tamura, Yoshimitsu Aoki, Yasuhiro Matsui, Kenji Iwata, Yutaka Satoh |
IECON | 1 |
| 2013 | Multiple players tracking and identification using group detection and player number recognition in sports videoabstractWe are interested in the problem of automatically tracking and identifying players in sports video. While there are many automatic multi-target tracking methods, in sports video, it is difficult to track multiple players due to frequent occlusions, quick motion of players and camera, and camera position. We propose tracking method that associates tracklets of a same player using results of player number recognition. To deal with frequent occlusions, we detect human region by level set method and then estimates if it is occluded group region or unoccluded individual one. Moreover, we associate tracklets using the results of player number recognition at each frame by keypoints-based matching with templates from multiple viewpoints, so that final tracklets include occluded region. Taiki Yamamoto, Hirokatsu Kataoka, Masaki Hayashi, Yoshimitsu Aoki, Kyoko Oshima, Masamoto Tanabiki |
IECON | 2 |
| 2012 | Extended CoHOG and particle filter by improved motion model for pedestrian active safetyabstractThe percentage of pedestrian deaths in traffic accidents is on the rise. In recent years, there have been calls for measures to be introduced to protect such vulnerable road users as pedestrians and cyclists. In this study, a method to detect pedestrians using an in-vehicle camera is presented. We improved the technology in detecting pedestrians with highly accurate images using a monocular camera. We were able to predict pedestrians' activities by monitoring them, and we developed an algorithm to recognize pedestrians and their movements more accurately. The effectiveness of the algorithm was tested using images taken on real roads. For the feature descriptor, we found that an extended co-occurrence histogram of oriented gradients, accumulating the integration of gradient intensities. In tracking step, we applied effective motion model using optical flow for Particle Filter tracking. These techniques are valified by using images captured on the real road. Hirokatsu Kataoka, Kimimasa Tamura, Yoshimitsu Aoki, Yasuhiro Matsui |
IECON | 1 |