Kenji Iwata

dblp:23/380 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
11since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 8 since 2021Systems, architecture and hardware · 6 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2024 The STVchrono Dataset: Towards Continuous Change Recognition in Time
abstract
Recognizing continuous changes offers valuable insights into past historical events, supports current trend analysis, and facilitates future planning. This knowledge is crucial for a variety of fields, such as meteorology and agriculture, environmental science, urban planning and construction, tourism, and cultural preservation. Currently available datasets in the field of scene change understanding primarily concentrate on two main tasks: the detection of changed regions within a scene and the linguistic description of the change content. Existing datasets focus on recognizing discrete changes, such as adding or deleting an object from two images, and largely rely on artificially generated images. Consequently, the existing change understanding methods primarily focus on identifying distinct object differences, overlooking the importance of continuous, gradual changes occurring over extended time intervals. To address the above issues, we propose a novel benchmark dataset, STVchrono, targeting the localization and description of long-term continuous changes in real-world scenes. The dataset consists of 71,900 photographs from Google Street View API taken over an 18-year span across 50 cities all over the world. Our STVchrono dataset is designed to support real-world continuous change recognition and description in both image pairs and extended image sequences, while also enabling the segmentation of changed regions. We conduct experiments to evaluate state-of-the- art methods on continuous change description and segmentation, as well as multimodal Large Language Models for describing changes. Our findings reveal that even the most advanced methods lag human performance, emphasizing the need to adapt them to continuously changing real-world scenarios. We hope that our benchmark dataset will further facilitate the research of temporal change recognition in a dynamic world. The STVchrono dataset is available at STVchrono Dataset.
Yanjun Sun, Yue Qiu 0001, Mariia Khan, Fumiya Matsuzawa, Kenji Iwata
CVPR5
2024 DailySTR: A Daily Human Activity Pattern Recognition Dataset for Spatio-temporal Reasoning
abstract
Recognizing daily human activities is essential for domestic robots to assist humans effectively in indoor environments. These activities typically involve sequences of interactions between humans and objects across different locations and times within a household. Identifying these events and understanding their temporal and spatial relationships is crucial for accurately modeling human behavior patterns. However, most current methods and datasets for human activity recognition focus on identifying singular events at specific moments and locations, neglecting the complexity of activities that span multiple times and places. To address this gap, we collected data on human activity patterns over a single day through crowdsourcing. Based on this, we introduce a novel synthetic video question-answering dataset. Our proposed dataset includes videos of daily activities accompanied by question-answer pairs that require models to reason about sequences of activities in both time and space. We evaluated state-of-the-art methods against our dataset, highlighting their limitations in handling the intricate spatio-temporal dynamics of human activity sequences. To improve upon these methods, we propose a two-stage model. The proposed model initially decodes the detailed content of individual videos using a transformer-based approach, then employs LLMs for advanced spatio-temporal reasoning across multiple videos. We hope our research provides valuable benchmarks and insights, paving the way for advancements in the recognition of daily human activity patterns.
Yue Qiu 0001, Shusaku Egami, Ken Fukuda, Natsuki Miyata, Takuma Yagi, Kensho Hara, Kenji Iwata, Ryusuke Sagawa
IROS7
2024 Subtle-Diff: A Dataset for Precise Recognition of Subtle Differences Among Visually Similar Objects
abstract
Visual inspection robots used in factories and outdoor environments require the ability to accurately recognize visual differences between similar objects and further verbalize the recognition results to present the differences to humans. Despite the application of Large Language Models (LLMs) and multimodal LLMs across various domains, our research highlights their insufficiency in verbalizing nuanced differences across images. To address this, we leveraged LLMs and image generation AI to develop a dataset aimed at assessing difference recognition capabilities. We introduced two novel tasks using this dataset: selecting images based on their visual differences and a conditional difference captioning task, and evaluated existing Vision-Language Models (VLMs) on these tasks. Our findings reveal that advanced models like GPT-4V can describe subtle differences with comparative expressions, yet they fall short of matching human performance across all attributes. This discrepancy between model and human recognition, especially in identifying easily discernible differences, suggests that most current models lack the ability to directly compare image pairs for difference detection. Consequently, we propose a new model that incorporates an image-text similarity approach in the difference recognition task, showing superior performance over existing models, including GPT-4V. Our dataset and findings will contribute to advancements in differencing objects and improve robotic applications in visual inspection and object picking. The dataset is available at DICTA challenge page.
Fumiya Matsuzawa, Yue Qiu 0001, Yanjun Sun, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh
IROS4
2023 Robust Recognition of Speaker Emotion With Difference Feature Extraction Using a Few Enrollment Utterances
abstract
This paper presents a novel approach to derive robust representations for speech emotion recognition by extracting speaker-independent features with the help of a speaker-dependent Gaussian mixture model (GMM). Since emotions are subjective and can vary greatly due to speaker behavior, incorporating speaker representations used for speaker verification tasks, such as x-vectors in past studies, have shown to improve the performance of speech emotion recognition, though they are not extracting speaker independent features. In this paper, we propose to derive embeddings that normalize speaker influence in the form of I-vectors, derived using a universal background model (UBM) trained only using the neutral emotion utterances from a single target speaker and an utterance-wise GMM also trained from the same speaker. We show through experiments on three datasets, that the proposed representations outperform methods that employ conventional x-vectors (which are not speaker-independent features) by approx. $3 \%$ absolute on average using as little as 4 enrollment utterances from the target speaker.
Daichi Hayakawa, Takehiko Kagoshima, Kenji Iwata, Norbert Braunschweiler, Rama Sanand Doddipatla
ASRU3
2023 Graph Representation for Order-aware Visual Transformation
abstract
This paper proposes a new visual reasoning formulation that aims at discovering changes between image pairs and their temporal orders. Recognizing scene dynamics and their chronological orders is a fundamental aspect of human cognition. The aforementioned abilities make it possible to follow step-by-step instructions, reason about and analyze events, recognize abnormal dynamics, and restore scenes to their previous states. However, it remains unclear how well current AI systems perform in these capabilities. Although a series of studies have focused on identifying and describing changes from image pairs, they mainly consider those changes that occur synchronously, thus neglecting potential orders within those changes. To address the above issue, we first propose a visual transformation graph structure for conveying order-aware changes. Then, we benchmarked previous methods on our newly generated dataset and identified the issues of existing methods for change order recognition. Finally, we show a significant improvement in order-aware change recognition by introducing a new model that explicitly associates different changes and then identifies changes and their orders in a graph representation.
Yue Qiu 0001, Yanjun Sun, Fumiya Matsuzawa, Kenji Iwata, Hirokatsu Kataoka
CVPR4
2023 Question Generation for Uncertainty Elimination in Referring Expressions in 3D Environments
abstract
We introduce a new task of question generation to eliminate the uncertainty of referring expressions in 3D indoor environments (3D-REQ). Referring to an object using natural language is one of the most common occurrences in daily human conversations; therefore, instructing robots to identify a certain object using natural language could be an essential task in var-ious robotic applications, such as room arrangement. However, human instructions are sometimes uncertain. Existing research on visual grounding using natural language in a 3D environment assumes that the referring expression can uniquely identify the object and does not consider that humans unconsciously give uncertain expressions. When faced with uncertainties, humans ask questions to gain further information. Inspired by the above observation, we propose a method that reduces uncertainty by asking questions when being given an obscure referring expression. The purpose of this method is to predict the positions of all candidate objects that satisfy the referring expressions in a 3D indoor environment and then to ask the appropriate questions to narrow down the target objects from them. To achieve this, we constructed a new 3D-REQ dataset, the input of which is a referring expression with uncertainties in the 3D environment and point clouds, and the output of which is the bounding boxes of all candidate objects satisfying the referring expression and a question to eliminate the uncertainty. To the best of our knowledge, 3D-REQ is the first effort to eliminate the uncertainty of referring expressions for object grounding in 3D environments.
Fumiya Matsuzawa, Yue Qiu 0001, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh
ICRA3
2023 VirtualHome Action Genome: A Simulated Spatio-Temporal Scene Graph Dataset with Consistent Relationship Labels
abstract
Spatio-temporal scene graph generation is an essential task in household activity recognition that aims to identify human-object interactions. Constructing a dataset with per-frame object region and consistent relationship annotations requires extremely high labor costs. Existing datasets sparsely annotate frames sampled from videos, resulting in the lack of dense spatio-temporal correlation in videos. Additionally, existing datasets contain inconsistent relationship annotations, leading to the problem of learning ambiguous temporal associations. Moreover, existing datasets mainly discuss relationships that can be inferred from a single frame, ignoring the significance of temporal associations. To resolve those issues, we created a simulated dataset with per-frame consistent annotations and introduced a range of relationships requiring both spatial and temporal context. Most existing methods explore spatial correlations within single images and do not explicitly consider the dynamic changes across frames. Therefore, we proposed a tracking-based approach that explicitly grasps spatio-temporal human-object interactions while simultaneously localizing humans and objects. Our proposed approach achieved state-of-the-art performance on scene graph generation and outperformed existing methods in scene graph localization by large margins on the proposed dataset. Moreover, the experiments show the efficacy of pre-training on the proposed dataset while adapting to a previous benchmark consisting of real daily videos, indicating the potential of the proposed dataset in real-world scenarios.
Yue Qiu 0001, Yoshiki Nagasaki, Kensho Hara, Hirokatsu Kataoka, Ryota Suzuki 0006, Kenji Iwata, Yutaka Satoh
WACV6
2023 3D Change Localization and Captioning from Dynamic Scans of Indoor Scenes
abstract
Daily indoor scenes often involve constant changes due to human activities. To recognize scene changes, existing change captioning methods focus on describing changes from two images of a scene. However, to accurately perceive and appropriately evaluate physical changes and then identify the geometry of changed objects, recognizing and localizing changes in 3D space is crucial. Therefore, we propose a task to explicitly localize changes in 3D bounding boxes from two point clouds and describe detailed scene changes, including change types, object attributes, and spatial locations. Moreover, we create a simulated dataset with various scenes, allowing generating data without labor costs. We further propose a framework that allows different 3D object detectors to be incorporated in the change detection process, after which captions are generated based on the correlations of different change regions. The proposed framework achieves promising results in both change detection and captioning. Furthermore, we also evaluated on data collected from real scenes. The experiments show that pretraining on the proposed dataset increases the change detection accuracy by +12.8% (mAP0.25) when applied to real-world data. We believe that our proposed dataset and discussion could provide both a new benchmark and in-sights for future studies in scene change understanding.
Yue Qiu 0001, Shintaro Yamamoto, Ryosuke Yamada, Ryota Suzuki 0006, Hirokatsu Kataoka, Kenji Iwata, Yutaka Satoh
WACV6
2022 Can Vision Transformers Learn without Natural Images?
abstract
Is it possible to complete Vision Transformer (ViT) pre-training without natural images and human-annotated labels? This question has become increasingly relevant in recent months because while current ViT pre-training tends to rely heavily on a large number of natural images and human-annotated labels, the recent use of natural images has resulted in problems related to privacy violation, inadequate fairness protection, and the need for labor-intensive annotations. In this paper, we experimentally verify that the results of formula-driven supervised learning (FDSL) framework are comparable with, and can even partially outperform, sophisticated self-supervised learning (SSL) methods like SimCLRv2 and MoCov2 without using any natural images in the pre-training phase. We also consider ways to reorganize FractalDB generation based on our tentative conclusion that there is room for configuration improvements in the iterated function system (IFS) parameter settings of such databases. Moreover, we show that while ViTs pre-trained without natural images produce visualizations that are somewhat different from ImageNet pre-trained ViTs, they can still interpret natural image datasets to a large extent. Finally, in experiments using the CIFAR-10 dataset, we show that our model achieved a performance rate of 97.8, which is comparable to the rate of 97.4 achieved with SimCLRv2 and 98.0 achieved with ImageNet.
Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata, Nakamasa Inoue, Yutaka Satoh
AAAI4
2022 GA-based Parameter Optimization of Image Processing for Contamination Inspection of Nonwoven Fabrics
abstract
The paper proposes the parameter optimization of image processing for contamination inspection of nonwoven fabrics. Currently, the automation of contamination inspection using image processing systems is being considered. In image processing, it is important to set the optimal parameters for the processing. However, it is necessary to search it from many combinations because there are some parameters. The proposed method searches for the optimal parameters based on a genetic algorithm. It reduces the search time in comparison with the conventional method. The paper indicates the effectiveness of the proposed method with the experimental results.
Nobuhiko Kumazawa, Sota Miyazaki, Yoshiyuki Hatta, Kazuaki Ito, Yukio Otsuka, Ryota Kitagawa, Kenji Iwata, Hidekazu Hirayu
IECON8
2021 Describing and Localizing Multiple Changes with Transformers
abstract
Change captioning tasks aim to detect changes in image pairs observed before and after a scene change and generate a natural language description of the changes. Existing change captioning studies have mainly focused on a single change. However, detecting and describing multiple changed parts in image pairs is essential for enhancing adaptability to complex scenarios. We solve the above issues from three aspects: (i) We propose a simulation-based multi-change captioning dataset; (ii) We benchmark existing state-of-the-art methods of single change captioning on multi-change captioning; (iii) We further propose Multi-Change Captioning transformers (MCCFormers) that identify change regions by densely correlating different regions in image pairs and dynamically determines the related change regions with words in sentences. The proposed method obtained the highest scores on four conventional change captioning evaluation metrics for multi-change captioning. Additionally, our proposed method can separate attention maps for each change and performs well with respect to change localization. Moreover, the proposed framework outperformed the previous state-of-the-art methods on an existing change captioning benchmark, CLEVR-Change, by a large margin (+6.1 on BLEU-4 and +9.7 on CIDEr scores), indicating its general ability in change captioning tasks. The code and dataset are available at the project page1.
Yue Qiu 0001, Shintaro Yamamoto, Kodai Nakashima, Ryota Suzuki 0006, Kenji Iwata, Hirokatsu Kataoka, Yutaka Satoh
ICCV5
2019 Slot Filling with Weighted Multi-Encoders for Out-of-Domain Values
Yuka Kobayashi, Takami Yoshida, Kenji Iwata, Hiroshi Fujimura
INTERSPEECH3
2018 Fashion Culture Database: Construction of Database for World-wide Fashion Analysis
abstract
The paper presents a novel concept that analyzes and visualizes worldwide fashion styles. Our goal is to reveal web-based viral fashion styles. To achieve the fashion-based analysis, we have collected fashion culture database (FCDB), which consists of 76 million geo-tagged images in 16 cosmopolitan cities. The database allows us to grasp a trend of mixed fashion styles with a fashion-based descriptor and codeword vector. In order to unveil web-based fashion trends in the FCDB, we applied a simple technique that is a temporal subtraction between consecutive codeword vectors in two different times. In the experiments, we show the analysis of fashion trends and fashion-based city similarity in a social media. As the result of large-scale data collection, we achieved world-level fashion visualization.
Kaori Abe, Munetaka Minoguchi, Teppei Suzuki, Naofumi Akimoto, Yue Qiu 0001, Ryota Suzuki 0006, Kenji Iwata, Yutaka Satoh, Hirokatsu Kataoka
ICARCV8
2018 Semantic Change Detection
abstract
Change detection is the study of detecting changes between two different images of a scene taken at different times. The change detection methodology can provide us information in which area images changed time by time. However, for application use, especially on disaster investigation, it is highly required to understand not only where but also what changes are occurred in high precision and resolution. The paper proposes the concept of semantic change detection, which involves intuitively inserting semantic meaning into detected change areas. We mainly focus on the novel semantic segmentation in addition to a conventional change detection approach. In order to solve this problem and obtain a high-level of performance, we propose an improvement to the hypercolumns representation, hereafter known as hypermaps, which effectively uses convolutional maps obtained from convolutional neural networks (CNNs). We also employ multi-scale feature representation captured by different image patches. We applied our method to the TSUNAMI panoramic change detection dataset (TSUNAMI dataset), and re-annotated the changed areas of the dataset via semantic classes. The results show that our multi-scale hypermaps provided outstanding performance on the re-annotated TSUNAMI dataset.
Munetaka Minoguchi, Ryota Suzuki 0006, Akio Nakamura, Kenji Iwata, Yutaka Satoh, Hirokatsu Kataoka
ICARCV5
2018 Occlusion Handling Human Detection with Refocused Images
abstract
The paper presents a novel robust human detection method based on camera array system to broaden the application range for human detection. Currently, even by using a deep neural network (DNN), it is difficult to detect a hardly occluded human. In the camera array system, we consider how to distinctly show a human occluded by an environmental condition. The generated refocused images by the camera array system allow us to remove the effect of the noises. Although refocused images have not been utilized in conventional human detection, we believe that the refocused images are beneficial for improving the detection performance, especially in severe conditions. To execute the experiments, we have collected Refocused Human DataBase (RHDB) with the camera array system. By using HOG+SVM with a monocular camera (at an almost random rate of 54.8%), the refocused images made the +10.1% improvement (64.9%) by noticeably showing a human. The combined representation of refocused images and AlexNet achieved 94.6% on the RHDB. Moreover, our final model recorded 98.0% with an attention-layer and fine-tuned parameters.
Hirokatsu Kataoka, Shuhei Ohki, Kenji Iwata, Yutaka Satoh
ICPR3
2018 Out-of-Domain Slot Value Detection for Spoken Dialogue Systems with Context Information
abstract
This paper proposes an approach to detecting-of-domain slot values from user utterances in spoken dialogue systems based on contexts. The approach detects keywords of slot values from utterances and consults domain knowledge (i.e., an ontology) to check whether the keywords are-of-domain. This can prevent the systems from responding improperly to user requests. We use a Recurrent Neural Network (RNN) encoder-decoder model and propose a method that uses only in-domain data. The method replaces word embedding vectors of the keywords corresponding to slot values with random vectors during training of the model. This allows using context information. The model is robust against over-fitting problems because it is independent of the slot values of the training data. Experiments show that the proposed method achieves a 65% gain in F1 score relative to a baseline model and a further 13 percentage points by combining with other methods.
Yuka Kobayashi, Takami Yoshida, Kenji Iwata, Hiroshi Fujimura, Masami Akamine
SLT3
2016 Recognition of Transitional Action for Short-Term Action Prediction using Discriminative Temporal CNN Feature
Hirokatsu Kataoka, Yudai Miyashita, Masaki Hayashi, Kenji Iwata, Yutaka Satoh
BMVC4
2015 Co-occurrence probability-based pixel pairs background model for robust object detection in dynamic scenes
Dong Liang 0008, Shun'ichi Kaneko, Manabu Hashimoto, Kenji Iwata, Xinyue Zhao
Pattern Recognit.4
2014 Extended Co-occurrence HOG with Dense Trajectories for Fine-Grained Activity Recognition
Hirokatsu Kataoka, Kiyoshi Hashimoto, Kenji Iwata, Yutaka Satoh, Nassir Navab, Slobodan Ilic, Yoshimitsu Aoki
ACCV (5)3
2013 Co-occurrence-based adaptive background model for robust object detection
abstract
An illumination-invariant background model for detecting objects in dynamic scenes is proposed. It is robust in the cases of sudden illumination fluctuation as well as burst moving background. Unlike previous works, it distinguishes objects from a dynamic background using co-occurrence character between a target pixel and its supporting pixels in the form of multiple pixel pairs. Experiments used several challenging datasets that proved the robust performance of object detection in various environments.
Dong Liang 0008, Shun'ichi Kaneko, Manabu Hashimoto, Kenji Iwata, Xinyue Zhao, Yutaka Satoh
AVSS4
2013 Robust feature descriptor and vehicle motion model with tracking-by-detection for active safety
abstract
The percentage of pedestrian deaths in traffic accidents is on the rise in Japan. In recent years, there have been calls for measures to be introduced to protect vulnerable road users such as pedestrians and cyclists. In this study, a method to detect and track pedestrians using an in-vehicle camera is presented to perform braking controls, warn the driver, and develop improved safety systems for pedestrians. We improved the technology of detecting pedestrians using highly accurate images obtained with a monocular camera. We were able to predict pedestrian activity by monitoring the images, and developed an algorithm with which to recognize pedestrians and their movements more accurately. The effectiveness of the algorithm was tested using images taken on real roads. For the feature descriptor, we used an extended co-occurrence histogram of oriented gradients (ECoHOG) that accumulated the integration of gradient intensities. In the tracking step, we applied an effective motion model using optical flow and the proposed feature descriptor ECoHOG in a tracking-by-detection framework. These techniques were verified using images captured on the real road.
Hirokatsu Kataoka, Kimimasa Tamura, Yoshimitsu Aoki, Yasuhiro Matsui, Kenji Iwata, Yutaka Satoh
IECON5
2011 Robust adapted object detection under complex environment
abstract
In this paper, we present a novel robust technique for background subtraction in different complex conditions (e.g. sudden illumination changes, swaying leaves, and camera vibrations). Unlike the previous works, the proposed method utilizes multiple point pairs that exhibit a stable statistical intensity relationship as a background model. The intensity difference between pixels of the pair is much more stable than the intensity of a single pixel, especially in varying environments. Furthermore, our proposed method focuses more on the history of global spatial correlations between pixels than on the history of any given pixel or local spatial correlations. we also adopt an adapted judgement criterion to ensure our method displays well in real-time detection. The approach has been compared with the state of the art on videos from several challenging datasets (PETS, Wallflower, and i-Lids), demonstrating that superior object detection is achieved.
Xinyue Zhao, Yutaka Satoh, Hidenori Takauji, Shun'ichi Kaneko, Kenji Iwata, Ryushi Ozaki
AVSS5
2011 Cancer detection from biopsy images using probabilistic and discriminative features
abstract
In the cancer detection from stained biopsy images, it is important to extract histologically discriminative characteristics. For this purpose, we propose a novel method to extract statistical and morphological features. At the first stage, we estimate cell component memberships at each pixel by applying an expectation maximization (EM) algorithm to the color information. Next we calculate the local co-occurrence of the memberships as image features. And then, linear discriminant analysis (LDA) is applied to those features for final decision of whether cancer or not, with enhancing the discrimination. In the experiments on real biopsy images of cancers, the resulting detection accuracy is superior to the other methods.
Atsushi Yaguchi, Takumi Kobayashi 0001, Kenji Watanabe, Kenji Iwata, Tadaaki Hosaka, Nobuyuki Otsu
ICIP4
2011 Object detection based on a robust and accurate statistical multi-point-pair model
Xinyue Zhao, Yutaka Satoh, Hidenori Takauji, Shun'ichi Kaneko, Kenji Iwata, Ryushi Ozaki
Pattern Recognit.5
2011 Semi-synchronous speech and pen input for mobile user interfaces
Koichi Shinoda, Yasushi Watanabe, Kenji Iwata, Ryuta Nakagawa, Sadaoki Furui
Speech Commun.3
2008 Robust spoken term detection using combination of phone-based and word-based recognition
abstract
We propose a robust spoken term detection method against word recognition errors using a combination of phone-based and word-based recognition.Conventional methods based on similar frameworks are problematic because phone-based recognition produces a large number of insertion errors.In our method, different substitution penalties are assigned for phone pairs to reduce such errors.We evaluated our method using the corpus of spontaneous Japanese.When recall was fixed at 50%, precision improved to 4.4 points above detection using only word-based recognition.We also report here on the effectiveness of optimization of the combination weight for each keyword.
Kenji Iwata, Koichi Shinoda, Sadaoki Furui
INTERSPEECH1
2007 Semi-Synchronous Speech and Pen Input
abstract
This paper proposes a new interface method using semi-synchronous speech and pen input for mobile environments. In this interface, a user speaks while writing, where pen input complements speech to achieve higher recognition performance than speech alone. A multimodal recognition algorithm that can handle the asynchronicity of the two modes using a segment-based unification scheme is proposed. This method is evaluated under noisy conditions with four different pen-input interfaces: character, stroke, pen-touch, and point-to-character, each of which is assumed to be given for a phrase unit in speech. It is confirmed that the recognition accuracy is improved by the proposed method in comparison with that by speech alone in all the pen-input conditions.
Yasushi Watanabe, Kenji Iwata, Ryuta Nakagawa, Koichi Shinoda, Sadaoki Furui
ICASSP (4)2
2007 Application of the Unusual Motion Detection Using CHLAC to the Video Surveillance
Kenji Iwata, Yutaka Satoh, Takumi Kobayashi 0001, Ikushi Yoda, Nobuyuki Otsu
ICONIP (2)1
2006 Hybrid Camera Surveillance System by Using Stereo Omni-directional System and Robust Human Detection
Kenji Iwata, Yutaka Satoh, Ikushi Yoda, Katsuhiko Sakaue
PSIVT1
2003 Robust Facial Parts Detection by Using Four Directional Features and Relaxation Matching
Kenji Iwata, Hitoshi Hongo, Kazuhiko Yamamoto, Yoshinori Niwa
KES1
2001 Book Cover Identification by Using Four Directional Features Filed for a Small-Scale Library System
abstract
This paper describes an identification method for book-cover images. It is used in a novel library system using images of books and the borrowers' faces. A returned book is identified from the lent books by processing the images of the book covers. A low-resolution field with four directional features is used to identify the books. These features can be compatible with computation time and performance. To identify a series of books which are similar, but different in just a small area, partial area matching is used. Experiments show the effectiveness of the proposed methods.
Kenji Iwata, Kazuhiko Yamamoto, Masakazu Yasuda, Kunihito Kato, Masahiko Ishida, Keishi Murata
ICDAR1
1988 Logic Simulation System Using Simulation Processor (SP)
Minoru Saitoh, Kenji Iwata, Akiko Nokamura, Makoto Kakegawa, Junichi Masuda, Hirofumi Hamamura, Fumiyasu Hirose, Nobuaki Kawato
DAC2