Hiroshi Murase

dblp:92/4484 · DBLP profile ↗
← Back
151ranked-venue papers
12as first author
17since 2021 · last 2026
0000-0002-8103-9294ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 88 · 12 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 85 · 5 first-author · 4 since 2021Systems, architecture and hardware · 7Databases, data management, data science and information retrieval · 7Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021
YearPublicationVenuePosition
2026 Resolving the Inherent Contextual Insufficiency in Referring Image Segmentation with Global Semantic Priors
Chong Yi, Jialei Chen 0001, Seigo Ito, Hiroshi Murase, Daisuke Deguchi
ICPR (6)4
2026 Semantic-Centric Alignment for Zero-shot Panoptic Segmentation with Limited Data
Jialei Chen 0001, Daisuke Deguchi, Xu Zheng 0002, Seigo Ito, Hiroshi Murase
Int. J. Comput. Vis.6
2026 Training-Free Open-Vocabulary Semantic Segmentation with Context Pyramid Refinement
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Hiroshi Murase, Daisuke Deguchi
Int. J. Comput. Vis.6
2026 Correction: Training-Free Open-Vocabulary Semantic Segmentation with Context Pyramid Refinement
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Hiroshi Murase, Daisuke Deguchi
Int. J. Comput. Vis.6
2026 CLIP-to-Seg Distillation for Zero-Shot Semantic Segmentation
Jialei Chen 0001, Zhenzhen Quan, Xu Zheng 0002, Daisuke Deguchi, Hiroshi Murase
IEEE Trans. Circuits Syst. Video Technol.6
2025 Semantic matters: A constrained approach for zero-shot video action recognition
abstract
Zero-shot video action recognition has advanced significantly due to the adaptation of visual-language models, such as CLIP, to video domains. However, existing methods attempt to adapt CLIP to video tasks by leveraging temporal information, neglecting the semantic information (i.e. the latent categories and their relationships) within videos. In this paper, we propose a Semantic Constrained CLIP (SC-CLIP) approach that leverages semantic information to adjust CLIP for video recognition while ensuring its performance on unseen data. SC-CLIP comprises a semantic-related query generation module and a semantic constrained cross attention module. First, the semantic-related query generation module clusters dense tokens from CLIP to generate semantic-related mask. The semantic-related query is then derived by pooling the adapted CLIP output using the semantic-related mask. Next, the semantic constrained cross attention module feeds the generated semantic-related query back into CLIP to probe semantic-related values, enhancing their ability to leverage the vision-language matching capabilities of CLIP. By generating semantic-related query, the semantic information aids in distinguishing similar actions, thereby improving performance on unseen samples. Experimental results on three zero-shot action recognition benchmarks show improvements of up to 1.9% and 2% in harmonic mean under two settings. Code is available at https://github.com/quanzhenzhen/SC-CLIP .
Zhenzhen Quan, Jialei Chen 0001, Daisuke Deguchi, Hiroshi Murase
Pattern Recognit.7
2024 Early Detection of At-risk Students Through Leaning-Activity Forecasting
abstract
With the widespread adoption of digital technologies such as digital textbooks, it has become feasible to collect daily logs of students' learning activities. Accordingly, there has been a growing trend in research using these logs. One of these areas is focusing on predicting grade of each student based on these learning activity logs. However, previous research focused on detecting At-risk students when learning activity logs for all lectures are available. This is not applicable for detection at the first few lectures (i.e. weeks) required in practical usage scenarios. We call this scenario as "early detection" in this paper. However, in early detection, the accuracy of at-risk detection tends to decrease. To solve this problem, we propose a Learning-activity Forecasting Network (LFNet) that improves the accuracy of early detection by aligning the embedding of the first few lectures with that of all lectures. Through experiments on learning activity logs of actual lectures, we confirmed that the proposed method could achieve high At-risk detection accuracy even from the first few lectures of learning activity logs.
Yuya Ozaki, Daisuke Deguchi, Haruya Kyutoku, Hiroshi Murase
ICCE4
2024 Frozen is better than learning: A new design of prototype-based classifier for semantic segmentation
abstract
Semantic segmentation models comprise an encoder to extract features and a classifier for prediction. However, the learning of the classifier suffers from the ambiguity which is caused by two factors: (1) the weights of a classifier for similar categories may have positive similarities lowing the performance for similar categories, named correlation ambiguity, and (2) the classifier is prone to predict the category with a larger ℓ2 norm and vice versa, termed prior ambiguity. To comedy the issues, we propose Category-Basis Prototype (CBP), frozen and mutually orthogonalized prototypes with equalℓ2 norm. Orthogonalization prevents the prototypes from being similar to each other and the equality decouples the prediction from the ℓ2 norm. To better shape the feature space, we propose Online Centroid Contrastive Loss (OCCL) equipped with centroid and category-level losses. Experiments show that our method yields compelling results over two widely applied benchmarks indicating the effectiveness of our methods.
Jialei Chen 0001, Daisuke Deguchi, Xu Zheng 0002, Hiroshi Murase
Pattern Recognit.5
2024 Texture-Guided Transfer Learning for Low-Quality Face Recognition
abstract
Although many advanced works have achieved significant progress for face recognition with deep learning and large-scale face datasets, low-quality face recognition remains a challenging problem in real-word applications, especially for unconstrained surveillance scenes. We propose a texture-guided (TG) transfer learning approach under the knowledge distillation scheme to improve low-quality face recognition performance. Unlike existing methods in which distillation loss is built on forward propagation; e.g., the output logits and intermediate features, in this study, the backward propagation gradient texture is used. More specifically, the gradient texture of low-quality images is forced to be aligned to that of its high-quality counterpart to reduce the feature discrepancy between the high- and low-quality images. Moreover, attention is introduced to derive a soft-attention (SA) version of transfer learning, termed as SA-TG, to focus on informative regions. Experiments on the benchmark low-quality face DB's TinyFace and QMUL-SurFace confirmed the superiority of the proposed method, especially more than 6.6% Rank1 accuracy improvement is achieved on TinyFace.
Meng Zhang 0042, Rujie Liu, Daisuke Deguchi, Hiroshi Murase
IEEE Trans. Image Process.4
2024 Toward Explainable End-to-End Driving Models via Simplified Objectification Constraints
abstract
The end-to-end driving models (E2EDMs) convert environmental information into driving actions using a complex transformation which makes E2EDMs have high prediction accuracy. Due to the black-box nature of transformation, the E2EDMs have low explainability. To solve this problem, explanation methods are used to generate explanations for observation. Based on current explanation methods, previous studies tried to further improve the explainability of E2EDMs by integrating an object detection module, however, these methods have many problems: Firstly, due to the requirement of the object detection module, they lack flexibility. Secondly, they neglect an essential property,i.e., simplicity, to improve explainability. In this paper, since humans prefer object-level and simple explanations in driving tasks, we argue that explainability is decided by two properties which are the objectification degree (the extent to which driving related-object features are utilized) and simplification degree (the simplicity of the explanation), thus we propose Simplified Objectification Branches (SOB) to improve the explainability of E2EDMs. Firstly, this structure could be integrated into any existing E2EDMs and thus have high flexibility. Secondly, the SOB explicitly improves the simplification degree without sacrificing the objectification degree of the explanations. By designing several indicators,i.e., heatmap satisfaction, driving action reproduction score, deception level,etc., we proved that SOB could help E2EDMs generate better explanations. Notably, the SOB could also further enhance E2EDMs’ prediction accuracy.
Daisuke Deguchi, Jialei Chen 0001, Hiroshi Murase
IEEE Trans. Intell. Transp. Syst.4
2023 Refined Objectification for Improving End-to-End Driving Model Explanation Persuasibility*
abstract
With the rapid development of deep learning, many end-to-end autonomous driving models with high prediction accuracy are developed. However, since autonomous driving technology is closely related to human life, users need to be convinced that end-to-end driving models (E2EDMs) not only have high prediction accuracy in known scenarios but also in practice for unknown scenarios. Therefore, engineers and end-users need to grasp the calculation methods of the E2EDMs based on the driving models’ explanations and ensure the explanations are satisfactory. However, few studies have focused on improving the explanation excellence.In this study, among many properties, we aim to improve the persuasibility of the explanation, we propose ROB (refined objectification branches), a structure that could be mounted to any type of existing E2EDMs. By persuasibility evaluation experiments, we demonstrate that one could improve the persuasibility of the explanations by mounting ROB to the E2EDM. As shown in Fig. 1, the focus area of the driving model accurately shrinks to the important and concise objects on account of ROB. We also perform an ablation study to further discuss each branch’s influence on persuasibility. In addition, we test ROB on multiple mainstream backbones and demonstrate that our structure could also improve the model’s prediction accuracy.
Daisuke Deguchi, Hiroshi Murase
IV3
2023 Implicit Interaction with an Autonomous Personal Mobility Vehicle: Relations of Pedestrians' Gaze Behavior with Situation Awareness and Perceived Risks
abstract
Interactions between pedestrians and autonomous personal mobility vehicle (APMV) will increase with the popularity of autonomous driving systems. However, when the APMVs are applied in a mixed traffic environment after manual driving PMV (MPMV) have been popular, pedestrians may feel unsafe in the interactions when they are uncertain about the driving intention of the APMV. This study seeks to find a surrogate measure for pedestrians’ understanding of driving intention and perceived safety during the interaction with an APMV. We conducted an experiment to measure the gaze duration and subjective evaluations of the participants when they interacted with a PMV in manual and autonomous driving modes. Pedestrians fixed their gaze at the APMV longer when they did not accurately understand the driving intention than when they understood it. Furthermore, the pedestrians perceived danger when they did not clearly understand the driving intention of the APMV. Besides, these factors were different when pedestrians interact with an MPMV and an APMV.
Hailong Liu 0001, Takatsugu Hirayama, Luis Yoichi Morales Saiki, Hiroshi Murase
Int. J. Hum. Comput. Interact.4
2022 Detection of Localization Failures Using Markov Random Fields With Fully Connected Latent Variables for Safe LiDAR-Based Automated Driving
abstract
Most of the recent automated driving systems assume the accurate functioning of localization. Unanticipated errors cause localization failures and result in failures in automated driving. An exact localization failure detection is necessary to ensure safety in automated driving; however, detection of the localization failures is challenging because sensor measurement is assumed to be independent of each other in the localization process. Owing to the assumption, the entire relation of the sensor measurement is ignored. Consequently, it is difficult to recognize the misalignment between the sensor measurement and the map when partial sensor measurement overlaps with the map. This paper proposes a method for the detection of localization failures using Markov random fields with fully connected latent variables. The full connection enables to take the entire relation into account and contributes to the exact misalignment recognition. Additionally, this paper presents localization failure probability calculation and efficient distance field representation methods. We evaluate the proposed method using two types of datasets. The first dataset is the SemanticKITTI dataset, whereby four methods are compared with the proposed method. The comparison results reveal that the proposed method achieves the most accurate failure detection. The second dataset is created based on log data acquired from the demonstrations that we conducted in Japanese public roads. The dataset includes several localization failure scenes. We apply the failure detection methods to the dataset and confirm that the proposed method achieves exact and immediate failure detection.
Naoki Akai, Yasuhiro Akagi, Takatsugu Hirayama, Takayuki Morikawa, Hiroshi Murase
IEEE Trans. Intell. Transp. Syst.5
2021 Persistent Homology in LiDAR-Based Ego-Vehicle Localization
abstract
Recently, various applications leveraging topological data analysis, in particular, persistent homology (PH), have been presented in many fields since PH provides a novel point cloud analysis method. In this work, we apply PH to LiDAR-based ego-vehicle localization applications. PH can extract translation and rotation invariant features from a point cloud. These features do not maintain local information of the point cloud, such as the edges and lines; however, they can abstract the global structure of the point cloud. A persistence image (PI) vectorizes the features and allows us to obtain fixed-size vectors despite the sizes of the source point clouds being different. Additionally, the size of the PI is not large even though the source point cloud is extremely big. We consider that these advantages are effective to loop closure detection, place categorization, and end-to-end global localization applications. Results reveal that it is difficult to improve the localization accuracy by simply applying PH owing to the basic concept of the topology that does not focus on exact shapes of the geometry. Therefore, we discuss how the advantages of PH can be utilized for the localization.
Naoki Akai, Takatsugu Hirayama, Hiroshi Murase
IV3
2021 Tell as You Imagine: Sentence Imageability-Aware Image Captioning
Kazuki Umemura, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
MMM (2)8
2021 Experimental stability analysis of neural networks in classification problems with confidence sets for persistence diagrams
Naoki Akai, Takatsugu Hirayama, Hiroshi Murase
Neural Networks3
2021 Soft-Boundary Label Relaxation with class placement constraints for semantic segmentation of the railway environment
Yuki Furitsu, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Hiroki Mukojima, Nozomi Nagamine
Pattern Recognit. Lett.5
2020 LFIR2Pose: Pose Estimation from an Extremely Low-resolution FIR image Sequence
abstract
In this paper, we propose a method for human pose estimation from a Low-resolution Far-InfraRed (LFIR) image sequence captured by a 16 × 16 FIR sensor array. Human body estimation from such a single LFIR image is a hard task. For training the estimation model, annotation of the human pose to the images is also a difficult task for human. Thus, we propose the LFIR2Pose model which accepts a sequence of LFIR images and outputs the human pose of the last frame, and also propose an automatic annotation system for the model training. Additionally, considering that the scale of human body motion is largely different among body parts, we also propose a loss function focusing on the difference. Through an experiment, we evaluated the human pose estimation accuracy with an original data set, and confirmed that human pose can be estimated accurately from an LFIR image sequence.
Saki Iwata, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa
ICPR5
2020 Ω-GAN: Object Manifold Embedding GAN for Image Generation by Disentangling Parameters into Pose and Shape Manifolds
abstract
In this paper, we propose Object Manifold Embedding GAN (Ω-GAN) to generate images of variously shaped and arbitrarily posed objects from a noise variable sampled from a distribution defined over the pose and the shape manifolds in a vector space. We introduce Parametric Manifold Sampling to sample noise variables from a distribution over the pose manifold to conditionally generate object images in arbitrary poses by tuning the pose parameter. We also introduce Object Identity Loss for clearly disentangling the pose and shape parameters, which allows us to maintain the shape of the object instance when only the pose parameter is changed. Through evaluation, we confirmed that the proposed Ω-GAN could generate variously shaped object images in arbitrary poses by changing the pose and shape parameters independently. We also introduce an application of the proposed method for object pose estimation, through which we confirmed that the object poses in the generated images are accurate.
Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ICPR4
2020 Median-Shape Representation Learning for Category-Level Object Pose Estimation in Cluttered Environments
abstract
In this paper, we propose an occlusion-robust pose estimation method of an unknown object instance in an object category from a depth image. In a cluttered environment, objects are often occluded mutually. For estimating the pose of an object in such a situation, a method that de-occludes the unobservable area of the object would be effective. However, there are two difficulties; occlusion causes the offset between the center of the actual object and its observable area, and different instances in a category may have different shapes. To cope with these difficulties, we propose a two-stage Encoder-Decoder model to extract features with objects whose centers are aligned to the image center. In the model, we also propose the Median-shape Reconstructor as the second stage to absorb shape variations in a category. By evaluating the method with both a large-scale virtual dataset and a real dataset, we confirmed the proposed method achieves good performance on pose estimation of an occluded object from a depth image.
Hiroki Tatemichi, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Ayako Amma, Hiroshi Murase
ICPR6
2020 Hybrid Localization using Model- and Learning-Based Methods: Fusion of Monte Carlo and E2E Localizations via Importance Sampling
abstract
This paper proposes a hybrid localization method that fuses Monte Carlo localization (MCL) and convolutional neural network (CNN)-based end-to-end (E2E) localization. MCL is based on particle filter and requires proposal distributions to sample the particles. The proposal distribution is generally predicted using a motion model. However, because the motion model cannot handle unanticipated errors, the predicted distribution is sometimes inaccurate. The use of other ideal proposal distributions, such as the measurement model, can improve robustness against such unanticipated errors. This technique is called importance sampling (IS). However, it is difficult to sample the particles from such ideal distributions because they are not represented in the closed form. Recent works have proved that CNNs with dropout layers represent the posterior distributions over their outputs conditioned on the inputs and the CNN predictions are equivalent to sampling the outputs from the posterior. Therefore, the proposed method utilizes a CNN to sample the particles and fuses them with MCL via IS. Consequently, the advantages of both MCL and E2E localization can be simultaneously leveraged while preventing their disadvantages. Experiments demonstrate that the proposed method can smoothly estimate the robot pose, similar to the model-based method, and quickly re-localize it from the failures, similar to the learning-based method.
Naoki Akai, Takatsugu Hirayama, Hiroshi Murase
ICRA3
2020 3D Monte Carlo Localization with Efficient Distance Field Representation for Automated Driving in Dynamic Environments
abstract
This paper presents a LiDAR-based 3D Monte Carlo localization (MCL) with an efficient distance field (DF) representation method. To implement 3D MCL, high computing capacity is required because the likelihood of many pose candidates, i.e., particles, must be calculated in real time by comparing sensor measurements and a map. Additionally, a large-scale map is needed for allocation to embedded computers since autonomous vehicles are required to navigate wide areas. These make it difficult for 3D MCL implementation. This paper first presents an efficient DF representation method while considering the 3D LiDAR-based localization characteristics. Because each DF voxel has the closest distance from occupied voxels, swift comparison of the sensor measurements and map can be achieved. Consequently, 3D MCL using the likelihood field model (LFM) can be executed in real time. Furthermore, this paper presents a method for improving the localization robustness to environmental changes without increasing memory and computational cost from that of the LFM-based MCL. Through experiments using the SemanticKITTI dataset, we show that the presented method can efficiently and robustly work in dynamic environments.
Naoki Akai, Takatsugu Hirayama, Hiroshi Murase
IV3
2020 Automatic Interaction Detection Between Vehicles and Vulnerable Road Users During Turning at an Intersection
abstract
Interaction detection between vehicles and vulnerable road users (e.g. pedestrians and cyclists) is important for e.g. safety control and autonomous driving. However, there are many challenges for automatically detecting interactions, such as the ambiguity of defining when interaction is required in dynamic traffic activities among different road users and the lack of labeled data for training a machine learning detector. To overcome the challenges, we introduce a way to define whether or not interaction is required in various traffic scenes and create a large real-world dataset from a very challenging intersection. A sequence-to-sequence method that uses the object information and motion information of the traffic scenes extracted by a state-of-the-art object detector and from optical flow, respectively, is proposed for automatic interaction detection. The proposed method generates a probability of interaction at each short interval (<; 0.1 s) that represents the changing of interaction along a sequence. We obtain a baseline model that differentiates no interaction from interaction on the basis of the location and road user type from the detected object information. Compared with the baseline model, the empirical results of the proposed method demonstrate very accurate predictions for vehicle turning sequences with varying length.
Hao Cheng 0008, Hailong Liu 0001, Fumito Shinmura, Naoki Akai, Hiroshi Murase, Takatsugu Hirayama
IV5
2020 Imageability Estimation using Visual and Language Features
abstract
Imageability is a concept from Psycholinguistics quantizing the human perception of words. However, existing datasets are created through subjective experiments and are thus very small. Therefore, methods to automatically estimate the imageability can be helpful. For an accurate automatic imageability estimation, we extend the idea of a psychological hypothesis called Dual-Coding Theory, that discusses the connection of our perception towards visual information and language information, and also focus on the relationship between the pronunciation of a word and its imageability. In this research, we propose a method to estimate imageability of words using both visual and language features extracted from corresponding data. For the estimation, we use visual features extracted from low- and high-level image features, and language features extracted from textual features and phonetic features of words. Evaluations show that our proposed method can estimate imageability more accurately than comparative methods, implying the contribution of each feature to the imageability.
Chihaya Matsuhira, Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Keisuke Doman, Daisuke Deguchi, Hiroshi Murase
ICMR8
2020 Browsing Visual Sentiment Datasets Using Psycholinguistic Groundings
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
MMM (2)6
2020 More-Natural Mimetic Words Generation for Fine-Grained Gait Description
Hirotaka Kato, Takatsugu Hirayama, Ichiro Ide, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase
MMM (2)7
2020 Estimating the imageability of words by mining visual characteristics from crawled image data
Marc A. Kastner 0001, Ichiro Ide, Frank Nack, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
Multim. Tools Appl.7
2019 Exemplar-Based Pseudo-Viewpoint Rotation for White-Cane User Recognition from a 2D Human Pose Sequence
abstract
In recent years, various facilities are equipped to support visually impaired people, but accidents caused by visual disabilities still occur. In this paper, to support the visually-impaired people in a public space, we aim to classify whether a pedestrian image sequence obtained by a surveillance camera is a white-cane user or not from the temporal transition of a human pose represented as 2D coordinates. However, since the appearance of the 2D pose varies largely depending on the viewpoint of the pose, it is difficult to classify them. So, in this paper, we propose a method to rotate the viewpoint of a pose from various pseudo-viewpoints based on a pair of 2D poses simultaneously observed and classify the sequence by multiple classifiers corresponding to each viewpoint. Viewpoint rotation makes it possible to obtain pseudo-poses seen from various pseudo-viewpoints, extract richer pose features, and recognize white-cane users more accurately. Through an experiment, we confirmed that the proposed method improves the recognition rate by 12% compared to the method not employing viewpoint rotation.
Naoki Nishida 0003, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Jun Piao
AVSS5
2019 Driving Behavior Modeling Based on Hidden Markov Models with Driver's Eye-Gaze Measurement and Ego-Vehicle Localization
abstract
This paper presents a comparison of driving behavior modeling methods based on hidden Markov models (HMMs) with driver's eye-gaze measurement and ego-vehicle localization. Original HMMs are sometimes insufficient to model real-world scenarios. To overcome these limitations, extended HMMs have been proposed, e.g., autoregressive input-output HMMs (AIOHMMs). This paper first details AIOHMMs and presents ways to use them for driving behavior modeling. We compare the performance for behavior modeling and maneuver discrimination for six types of HMMs. The driving data for this work was gathered in our university campus with a car-like vehicle. Experimental results suggest that the hidden states can properly represent the average of the driving actions when the driving behaviors are accurately modeled by the HMMs. It is also suggested that surrounding and past information can be used to flexibly model the relationship between driving actions and related information.
Naoki Akai, Takatsugu Hirayama, Luis Yoichi Morales Saiki, Yasuhiro Akagi, Hailong Liu 0001, Hiroshi Murase
IV6
2019 Estimating the visual variety of concepts by referring to Web popularity
Marc A. Kastner 0001, Ichiro Ide, Yasutomo Kawanishi, Takatsugu Hirayama, Daisuke Deguchi, Hiroshi Murase
Multim. Tools Appl.6
2018 Which Content in a Booklet is he/she Reading? Reading Content Estimation using an Indoor Surveillance Camera
abstract
In this paper, we propose a method for estimating reading content in a booklet using an image captured by an indoor surveillance camera. Here, we assume that a reading content can be specified by estimating followings; what booklet, which page of the booklet, and which region in the page. We propose a reading booklet/page estimation method based on image search, and a reading region estimation method focusing on the body pose of the reader. We evaluated the method as a 44 classes classification problem, which consists of eleven pages of booklets and four regions in each pages. We achieved 25.6% in accuracy of the reading content estimation.
Yasutomo Kawanishi, Hiroshi Murase, Kazuyuki Tasaka, Hiromasa Yanagihara
ICPR2
2018 Mobile Robot Localization Considering Class of Sensor Observations
abstract
Localization robustness against environment dynamics is significant for robots to achieve autonomous navigation in unmodified environments. A basic method of improving the robustness of a robot is considering the sensor observations obtained from mapped obstacles and using them for localizing the robot's pose. This study proposes an observation model that considers the class of sensor observations, where “class” categorizes the sensor observations as those obtained from mapped and unmapped obstacles. In the proposed approach, the robot's pose and the class are estimated simultaneously. As a result, the robot's pose can be localized using the sensor observations obtained only from mapped obstacles. First, we evaluated the performance of the proposed approach using simulations. Further, we tested the proposed approach in a real-world mobile robot navigation competition, called “Tsukuba Challenge,” held in Japan. The robustness and effectiveness of the proposed approach against environment dynamics were verified from the experimental results.
Naoki Akai, Luis Yoichi Morales Saiki, Hiroshi Murase
IROS3
2018 Personal Mobility Vehicle Autonomous Navigation Through Pedestrian Flow: A Data Driven Approach for Parameter Extraction
abstract
In this paper we present a data driven approach for safe and smooth autonomous navigation of a personal mobility vehicle (PMV) when facing moving obstacles such as people and bicycles in public pedestrian paths. In a period of three months, data from five different persons driving the robotic PMV in an outdoor environment while facing pedestrians were collected. 2465 clean tracks around the vehicle together with PMVs trajectories were collected. We performed an analysis of the parameters involved for human-driven smooth navigation. Relevant parameters regarding PMV-Human interaction included distance to moving objects, passing side and velocities. Moreover, data suggests the existence of a social navigational distance for the PWv. For autonomous navigation we implemented a Frenet planner to achieve safe and smooth navigation for the passenger and pedestrians around. Experimental results in real pedestrian paths show that the PMV is capable of smoothly following its path while facing pedestrians and bicycles.
Luis Yoichi Morales Saiki, Naoki Akai, Hiroshi Murase
IROS3
2018 Gaze-Inspired Learning for Estimating the Attractiveness of a Food Photo
abstract
The number of food photos posted to the Web has been increasing. Most of the users prefer to post delicious-looking food photos. They, however, do not always look delicious. A previous work proposed a method for estimating the attractiveness of food photos, that is, the degree of how much a food photo looks delicious, as an assistive technology for taking a delicious-looking food photo. This method extracted image features from the entire food photo to evaluate the impression. In our work, we conduct a preference experiment where subjects are asked to compare a pair of food photos and measure their gaze. The proposed method extracts image features from local regions selected based on the gaze information and estimates the attractiveness of a food photo by learning regression parameters. Experimental results showed the effectiveness of extracting image features from outside the gaze regions rather than inside them.
Akinori Sato, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ISM7
2018 Reliability Estimation of Vehicle Localization Result
abstract
This paper proposes a method for estimation of the reliability of vehicle localization results. We previously proposed a fault detection method for indoor mobile robots using a convolutional neural network (CNN). Because image data is generally fed to a CNN, we feed image data obtained from the robot pose, occupancy grid map, and laser scan data to the CNN, which decides of whether localization has failed. The previous method also employed a Rao-Blackwellized particle filter to estimate the robot pose and reliability of this estimation simultaneously. However, it was difficult for vehicle robots to use the previous method as creating and processing image data is not a light computation process. In this study, we extend the previous method by improving the data fed to the CNN, thus making it possible for vehicle robots to perform simultaneous localization and estimation. This paper describes in detail the simultaneous estimation and shows that the reliability can be used as an exact criterion for detecting localization failures. Keywords-Vehicle Localization, Reliability
Naoki Akai, Luis Yoichi Morales Saiki, Hiroshi Murase
Intelligent Vehicles Symposium3
2018 Sparse Coding of Weather and Illuminations for ADAS and Autonomous Driving
abstract
Weather and illumination are critical factors in vision tasks such as road detection, vehicle recognition, and active lighting for autonomous vehicles and ADAS. Understanding the weather and illumination type in a vehicle driving view can guide visual sensing, control vehicle headlight and speed, etc. This paper uses sparse coding technique to identify weather types in driving video, given a set of bases from video samples covering a full spectrum of weather and illumination conditions. We sample traffic and architecture insensitive regions in each video frame for features and obtain clusters of weather and illuminations via unsupervised learning. Then, a set of keys are selected carefully according to the visual appearance of road and sky. For video input, sparse coding of each frame is calculated for representing the vehicle view robustly under a specific illumination. The linear combination of the basis from keys results in weather types for road recognition, active lighting, intelligent vehicle control, etc.
Jiang Yu Zheng, Hiroshi Murase
Intelligent Vehicles Symposium3
2018 Voting-based Hand-Waving Gesture Spotting from a Low-Resolution Far-Infrared Image Sequence
abstract
We propose a temporal spotting method of a hand gesture from a low-resolution far-infrared image sequence captured by a far-infrared sensor array. The sensor array captures the spatial distribution of far-infrared intensity as a thermal image by detecting far-infrared waves emitted from heat sources. It is difficult to spot a hand gesture from a sequence of thermal images captured by the sensor due to its low-resolution, heavy noise, and varying duration of the gesture. Therefore, we introduce a voting-based approach to spot the gesture with template matching-based gesture recognition. We confirm the effectiveness of the proposed temporal spotting method in several settings.
Yasutomo Kawanishi, Chisato Toriyama, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Tomoyoshi Aizawa, Masato Kawade
VCIP6
2017 Action recognition from extremely low-resolution thermal image sequence
abstract
This paper proposes a Deep Learning-based action recognition method from an extremely low-resolution thermal image sequence. The method recognizes daily actions by humans (e.g. walking, sitting down, standing up, etc.) and abnormal actions (e.g. falling down) without privacy concerns. While privacy concerns can be ignored, it is difficult to compute feature points and to obtain a clear edge of the human body from an extremely low-resolution thermal image. To address these problems, this paper proposes a Deep Learning-based action recognition method that combines convolution layers and an LSTM layer for learning spatio-temporal representation, whose inputs are the thermal images and their frame differences cropped by the gravity center of human regions. The effectiveness of the proposed method was confirmed through experiments.
Takayuki Kawashima, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Daisuke Deguchi, Tomoyoshi Aizawa, Masato Kawade
AVSS4
2017 Automatic Selection of Web Contents Towards Automatic Authoring of a Video Biography
abstract
In this paper, we propose a method for image selection using Web image search for automatic video biography authoring. In the proposed method, images are selected from the image search results considering their visual contents for inclusion in the video biography. Through evaluation, we confirmed the effectiveness of the proposed image selection method compared to a baseline method which simply selects the top 1 search result.
Ichiro Ide, Yasutomo Kawanishi, Kyoka Kunishiro, Frank Nack, Daisuke Deguchi, Hiroshi Murase
ISM6
2017 Summarization of News Videos Considering the Consistency of Auditory and Visual Contents
abstract
Since news videos are valuable sources of multimedia information on real-world events, there is a demand for viewing them efficiently. However, there is a problem that summarization methods based on auditory contents do not take into account the visual contents. In the case of news videos, due to its presentation style where audio contents and visual contents do not necessarily come from the same source, this could severely decrease the amount of informative visual contents included in the generated summarized video. Thus, we propose a method for summarizing a sequence of news videos considering the consistency of both auditory and visual contents. The proposed method first selects key-sentences from the auditory contents (Closed Caption) of each news story in the sequence, and then selects a shot within the news story whose "Visual Concepts" detected from the visual contents are the most consistent with the key-phrase. Finally, the audio segment corresponding to each key-phrase is overlapped onto the selected shot, and then concatenated to generate a summarized video. The effectiveness of the proposed method was confirmed on several news topics through a subjective experiment.
Ichiro Ide, Ryunosuke Tanishige, Keisuke Doman, Yasutomo Kawanishi, Daisuke Deguchi, Hiroshi Murase
ISM7
2017 Monocular localization within sparse voxel maps
abstract
We introduce a method that uses a single camera to localize a vehicle within a pre-constructed map consisting of a voxel occupancy grid and road-line marker positions. Sophisticated mapping hardware is capable of creating high-accuracy 3D maps of road environments, but localizing a vehicle within such maps is one of the challenges at the forefront of automated driving. A solution which is robust to dynamic environments, while using only inexpensive sensors, is a difficult problem. In addition, maps that enable precise localization consume a lot of data which is impractical for the expansive environments encountered in real-world road networks. We show how using the area of edge regions shared between rendered views of a compact voxel map and in-vehicle camera images can be coupled with non-linear optimization methods to determine the camera position and pose.
David Wong 0002, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium5
2017 Proposal of a spectral random dots marker using local feature for posture estimation
abstract
We propose a novel marker for robot's grasping task which has the following three aspects: (i) it is easy-to-find in a cluttered background, (ii) it is calculable for its posture (iii) its size is compact. The proposed marker is composed of a random dots pattern, and uses keypoint detection and a scale estimation by Spectral SIFT for dots detection and data decoding. The data is encoded by the scale size of dots, and the same dots in the marker work for both marker detection and data decoding. As a result, the proposed marker size can be compact. We confirmed the effectiveness of the proposed marker through experiments.
Norimasa Kobori, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
VR4
2017 Regression of feature scale tracklets for decimeter visual localization
David Wong 0002, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase
Image Vis. Comput.5
2016 A classification method of cooking operations based on eye movement patterns
abstract
We are developing a cooking support system that coaches beginners. In this work, we focus on eye movement patterns while cooking meals because gaze dynamics include important information for understanding human behavior. The system first needs to classify typical cooking operations. In this paper, we propose a gaze-based classification method and evaluate whether or not the eye movement patterns have a potential to classify the cooking operations. We improve the conventional N-gram model of eye movement patterns, which was designed to be applied for recognition of office work. Conventionally, only relative movement from the previous frame was used as a feature. However, since in cooking, users pay attention to cooking ingredients and equipments, we consider fixation as a component of the N-gram. We also consider eye blinks, which is related to the cognitive state. Compared to the conventional method, instead of focusing on statistical features, we consider the ordinal relations of fixation, blink, and the relative movement. The proposed method estimates the likelihood of the cooking operations by Support Vector Regression (SVR) using frequency histograms of N-grams as explanatory variables.
Hiroya Inoue, Takatsugu Hirayama, Keisuke Doman, Yasutomo Kawanishi, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ETRA7
2016 Moving camera background-subtraction for obstacle detection on railway tracks
abstract
We propose a method for detecting obstacles by comparing input and reference train frontal view camera images. In the field of obstacle detection, most methods employ a machine learning approach, so they can only detect pre-trained classes, such as pedestrian, bicycle, etc. This means that obstacles of unknown classes cannot be detected. To overcome this problem, we propose a background subtraction method that can be applied to moving cameras. First, the proposed method computes frame-by-frame correspondences between the current and the reference (database) image sequences. Then, obstacles are detected by applying image subtraction to corresponding frames. To confirm the effectiveness of the proposed method, we conducted an experiment using several image sequences captured on an experimental track. Its results showed that the proposed method could detect various obstacles accurately and effectively.
Hiroki Mukojima, Daisuke Deguchi, Yasutomo Kawanishi, Ichiro Ide, Hiroshi Murase, Masato Ukai, Nozomi Nagamine, Ryuta Nakasone
ICIP5
2016 Misclassification tolerable learning for robust pedestrian orientation classification
abstract
In this paper, we propose a multiclass classifier training method which reduces “fatal” misclassifications by cost-relaxation of “tolerable” misclassifications in one-against-all classifiers training, named misclassification tolerable learning. In a binary classifier in the one-against-all classifiers, we introduce a new class group “conceptually similar classes,” whose class labels are similar to the positive class. In the case of pedestrian orientation classification, the conceptually similar classes are defined as neighboring orientations to the positive orientation. We consider the misclassification of the conceptually similar classes to the positive class as tolerable misclassification. By relaxing the cost of the tolerable misclassifications, our proposed classification method reduces fatal misclassifications of non-similar classes. We evaluated the cost-relaxation effectiveness on several public datasets and confirmed that the proposed method outperforms the normal SVM on all of the datasets in the soft criterion by achieving 78.63% recognition rate on PDC Dataset.
Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi
ICPR4
2016 Robust pedestrian attribute recognition for an unbalanced dataset using mini-batch training with rarity rate
abstract
Pedestrian attributes are significant information for Advanced Driver Assistance System(ADAS). Pedestrian attributes such as body poses, face orientations and open umbrella are meant action or state of pedestrian. In general, this information is recognized using independent classifiers for each task. Performing all of these separate tasks is too time-consuming at the testing stage. In addition, the processing time increases with the number of tasks. To address this problem, multi-task learning or heterogeneous learning is able to train a single classifier to perform multiple tasks. In particular, heterogeneous learning is able to simultaneously train regression and recognition tasks, because reducing both training and testing time. However, heterogeneous learning tends to result in a lower accuracy rate for classes with a few training samples. In this paper, we propose a method to improve the performance of heterogeneous learning for such classes. We introduce a rarity rate based on the importance and class probability of each task. The appropriate rarity rate is assigned to each training sample. Thus, the samples in a mini-batch for training a deep convolutional neural network are augmented by this rarity rate to focus on the class with a few samples. Our heterogeneous learning approach with the rarity rate attains better performance on pedestrian attribute recognition, especially for classes representing open umbrellas.
Hiroshi Fukui, Takayoshi Yamashita, Yuji Yamauchi, Hironobu Fujiyoshi, Hiroshi Murase
Intelligent Vehicles Symposium5
2016 Parts Selective DPM for detection of pedestrians possessing an umbrella
abstract
In recent years, pedestrian detection from an in-vehicle camera has been attracting attention. However, in the case of a raining situation, the detection accuracy decreases because the head of a pedestrian tends to be occluded by an umbrella. In oder to handle such cases, in this paper, as a variation of the Deformable Part Model (DPM) which is widely used in the field of object recognition, we propose “Parts Selective DPM (PS-DPM)” which selectively chooses the original part filters and additional part filters trained independently. In the detection of pedestrians possessing an umbrella, the selection of head and umbrella parts will make pedestrian detection more robust to the occlusion. We conducted experiments to evaluate the performance of the proposed method. As a result, pedestrian detection with the proposed PS-DPM achieved high detection accuracy in rainy weather, compared with the detection by the conventional DPM. Moreover, we confirmed that it did not decrease the pedestrian detection accuracy in fine weather.
Yuto Shimbo, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium5
2015 Fast 3D edge detection by using decision tree from depth image
abstract
T3D edge detection from a depth image is an important technique of 3D object recognition in preprocessing. There are three types of 3D edges in a depth image called jump, convex roof, and concave roof edges. Conventional 3D edge detection based on ring operators has been proposed. The conventional ring operator can detect three types of 3D edges by classifying the response of Fourier transforms. Since the conventional method needs to apply Fourier transforms to all pixels of a depth image, real-time processing cannot be done due to high computational cost. Therefore, this paper presents a fast and reliable method of detecting three types of 3D edges by using a decision tree. The decision tree is trained under supervised learning from numerous synthesized depth images and labels by capturing depth relations between candidate pixels and pixels on a ring operator to classify 3D edges. The experimental results revealed that the proposed method has 25 times faster than the conventional method. This paper also presents some examples of 3D line and 3D convex corner detection based on results obtained with the proposed method.
Masaya Kaneko, Takahiro Hasegawa, Yuji Yamauchi, Takayoshi Yamashita, Hironobu Fujiyoshi, Hiroshi Murase
IROS6
2015 Pedestrian detection based on deep convolutional neural network with ensemble inference network
abstract
Pedestrian detection is an active research topic for driving assistance systems. To install pedestrian detection in a regular vehicle, however, there is a need to reduce its cost and ensure high accuracy. Although many approaches have been developed, vision-based methods of pedestrian detection are best suited to these requirements. In this paper, we propose the methods based on Convolutional Neural Networks (CNN) that achieves high accuracy in various fields. To achieve such generalization, our CNN-based method introduces Random Dropout and Ensemble Inference Network (EIN) to the training and classification processes, respectively. Random Dropout selects units that have a flexible rate, instead of the fixed rate in conventional Dropout. EIN constructs multiple networks that have different structures in fully connected layers. The proposed methods achieves comparable performance to state-of-the-art methods, even though the structure of the proposed methods are considerably simpler.
Hiroshi Fukui, Takayoshi Yamashita, Yuji Yamauchi, Hironobu Fujiyoshi, Hiroshi Murase
Intelligent Vehicles Symposium5
2015 Pedestrian orientation classification utilizing single-chip coaxial RGB-ToF camera
abstract
This paper proposes a method for pedestrian orientation classification. In image recognition, the accuracy is often degraded by the influence of background. In addition, it is also difficult to remove the background and extract only the human body from an image. To overcome these problems, we utilize a single-chip RGB-ToF camera. This camera can acquire RGB and depth images along the same optical axis at the same moment, and thus segmentation of the RGB image becomes easier by using the coaxial depth image. Our proposed method segmented a human body from its background accurately, which lead to the improvement of the accuracy of pedestrian orientation classification.
Fumito Shinmura, Yasutomo Kawanishi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Hironobu Fujiyoshi
Intelligent Vehicles Symposium5
2014 Spatial People Density Estimation from Multiple Viewpoints by Memory Based Regression
abstract
Crowd analysis using cameras has attracted much attention for public safety and marketing. Among techniques of the crowd analysis, we focus on spatial people density estimation which estimates the number of people for each small area in a floor region. However, spatial people density cannot be estimated accurately for an area far from the camera because of the occlusion by people in a closer area. Therefore, we propose a method using a memory based regression method with images captured from cameras from multiple viewpoints. This method is realized by looking up a table that consists of correspondences between people density maps and crowd appearances. Since the crowd appearances include situations where various occlusions occur, an estimation robust to occlusion should be realized. In an experiment, we examined the effectiveness of the proposed method.
Yoshimune Tabuchi, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takayuki Kurozumi, Kunio Kashino
ICPR5
2014 Scene Duplicate Detection from News Videos Using Image-Audio Matching Focusing on Human Faces
abstract
As one tool for structuring a massive volume of archived news videos based on their semantic contents, this paper proposes a method to detect scene duplicates from news videos. A scene duplicate is a pair of video segments taken at the same event from different viewpoints. Referring to the audio channel is effective to detect scene duplicates regardless of viewpoints, but it cannot be relied on when external audio sources (e.g. Narrations, sound effects) overlap the original one. In contrast, the image channel can be useful in most cases, although significant difference in viewpoints affect the detection. The proposed method integrates the information from these two channels in order to improve the accuracy of scene duplicate detection from news videos. The performance of the proposed method was evaluated through an experiment with actual broadcast news videos. As a result, we obtained the higher detection accuracies in both recall and precision. Therefore, we confirmed the effectiveness of the proposed method.
Haruka Kumagai, Keisuke Doman, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ISM5
2014 Estimation of traffic sign visibility considering local and global features in a driving environment
abstract
This paper proposes a camera-based visibility estimation method for a traffic sign. The visibility here indicates how a visual target is easy to be detected and recognized by a human driver (not a machine). This research aims at realizing a nuisance-free driver assistance system which sorts out information depending on the visibility of a visual target, in order to prevent driver distraction. Our previous study on estimating the visibility of a traffic sign considered only the effect of the local region around a target, assuming the situation that a driver's gaze is around it. The proposed method integrates both the local features and global features in a driving environment without such an assumption. The global features evaluate the positional relationships between traffic signs and the appearance around the fixation point of a driver's gaze, which considers the effect of the driver's entire field of view. Experimental results showed the effectiveness of incorporating the global features for estimating the visibility of a traffic sign.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Utsushi Sakai
Intelligent Vehicles Symposium6
2014 Single camera vehicle localization using SURF scale and dynamic time warping
abstract
Vehicle ego-localization is an essential process for many driver assistance and autonomous driving systems. The traditional solution of GPS localization is often unreliable in urban environments where tall buildings can cause shadowing of the satellite signal and multipath propagation. Typical visual feature based localization methods rely on calculation of the fundamental matrix which can be unstable when the baseline is small. In this paper we propose a novel method which uses the scale of matched SURF image features and Dynamic Time Warping to perform stable localization. By comparing SURF feature scales between input images and a pre-constructed database, stable localization is achieved without the need to calculate the fundamental matrix. In addition, 3D information is added to the database feature points in order to perform lateral localization, and therefore lane recognition. From experimental data captured from real traffic environments, we show how the proposed system can provide high localization accuracy relative to an image database, and can also perform lateral localization to recognize the vehicle's current lane.
David Wong 0002, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium4
2014 Estimation of the Representative Story Transition in a Chronological Semantic Structure of News Topics
abstract
It is important to track the flow of topics to thoroughly understand the contents. Accordingly, a method that structures the chronological semantic relations between news stories, namely a "topic thread structure" has been proposed. It allows the comprehensive understanding of a topic by chronologically tracking stories one by one from the initial story. However, this task imposes a user to watch many stories when it contains various sub-topics. Thus, we propose a method that estimates the representative story transition in a topic thread structure. In the proposed method, features obtained from a story and those from the topic thread structure are used for the estimation. We confirmed the effectiveness of the proposed method by comparing the results obtained from the proposed method to the ground truth obtained from votes in a subjective experiment.
Kosuke Kato, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ICMR4
2014 Event Detection based on Twitter Enthusiasm Degree for Generating a Sports Highlight Video
abstract
This paper presents a Twitter-based event detection method based on "Twitter Enthusiasm Degrees (TED)" toward generating a highlight video of a sports game. Existing methods not only depend on both languages and sports types but also often falsely detect non-target events. In contrast, the proposed method detects sports events using TEDs calculated from several kinds of string features independent of languages and sports. We applied the proposed method to actual sports games, and compared the detected events with the events present in broadcasted highlight videos, and confirmed the effectiveness and the language and sports type independencies of the proposed method.
Keisuke Doman, Taishi Tomita, Ichiro Ide, Daisuke Deguchi, Hiroshi Murase
ACM Multimedia5
2013 Pedestrian detection by scene dependent classifiers with generative learning
abstract
Recently, pedestrian detection from in-vehicle camera images is becoming an crucial technology for Intelligent Transportation Systems (ITS). However, it is difficult to detect pedestrians accurately in various scenes by obtaining training samples. To tackle this problem, we propose a method to construct scene dependent classifiers to improve the accuracy of pedestrian detection. The proposed method selects an appropriate classifier based on the scene information that is a category of appearance associated with location information. To construct scene dependent classifiers, the proposed method introduces generative learning for synthesizing scene dependent training samples. Experimental results showed that the detection accuracy of the proposed method outperformed the comparative method, and we confirmed that scene dependent classifiers improved the accuracy of pedestrian detection.
Hidefumi Yoshida, Daichi Suzuo, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Takashi Machida, Yoshiko Kojima
Intelligent Vehicles Symposium5
2013 Detection of Biased Broadcast Sports Video Highlights by Attribute-Based Tweets Analysis
Takashi Kobayashi 0001, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
MMM (2)5
2012 Robust Face Super-Resolution Using Free-Form Deformations for Low-Quality Surveillance Video
abstract
Recently, the demand for face recognition to identify persons from surveillance video cameras has rapidly increased. Since surveillance cameras are usually placed at positions far from a person's face, the quality of face images captured by the cameras tends to be low. This degrades the recognition accuracy. Therefore, aiming to improve the accuracy of the low-resolution-face recognition, we propose a video-based super-resolution method. The proposed method can generate a high-resolution face image from low-resolution video frames including non-rigid deformations caused by changes of face poses and expressions without using any positional information of facial feature points. Most existing techniques use the facial feature points for image alignment between the video frames. However, it is difficult to obtain the accurate positions of the feature points from low-resolution face images. To achieve the alignment, the proposed method uses a free-form deformation method that flexibly aligns each local region between the images. This enables super-resolution of face images from low-resolution videos. Experimental results demonstrated that the proposed method improved the performance of super-resolution for actual videos in terms of both image quality and face recognition accuracy.
Tomonari Yoshida, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ICME5
2012 Estimation of the human performance for pedestrian detectability based on visual search and motion features
Masashi Wakayama, Daisuke Deguchi, Keisuke Doman, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
ICPR5
2012 Smart VideoCooKing: a multimedia cooking recipe browsing application on portable devices
abstract
This demo presents "Smart VideoCooKing" which is a multimedia cooking recipe browsing application on portable Android devices. A multimedia cooking recipe is a cooking recipe where each cooking operation is associated with a corresponding video clip describing it, aimed to facilitate the understanding of cooking operations. In combination with third-party applications, "Smart VideoCooKing" provides useful functions such as playing cooking video clips describing cooking operations quickly, searching information of ingredients easily, and reading aloud cooking directions.
Keisuke Doman, Cheng Ying Kuai, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ACM Multimedia5
2011 Low Resolution QR-Code Recognition by Applying Super-Resolution Using the Property of QR-Codes
abstract
This paper proposes a method for low resolution QR-code recognition. A QR-code is a two-dimensional binary symbol that can embed various information such as characters and numbers. To recognize a QR-code correctly and stably, the resolution of an input image should be high. In practice, however, recognition of a QR-code is usually difficult due to low resolution when it is captured from a distance. In this paper, we propose a method to improve the performance of low resolution QR-code recognition by using the super-resolution technique that generates a high resolution image from multiple low-resolution images. Although a QR-code is a binary pattern, it is observed as a grayscale image due to the degradation through the capturing process. Especially the pixels around the borders between white and black regions become ambiguous. To overcome this problem, the proposed method introduces a binary pattern constraint to generate super-resolved images appropriate for recognition. Experimental results showed that a recognition rate of 98% can be achieved by the proposed method, which is a 15.7% improvement in comparison with a method using a conventional super-resolution method.
Yuji Kato, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICDAR5
2011 Detection of Inconsistency Between Subject and Speaker Based on the Co-occurrence of Lip Motion and Voice Towards Speech Scene Extraction from News Videos
abstract
We propose a method to detect the inconsistency between a subject and the speaker for extracting speech scenes from news videos. Speech scenes in news videos contain a wealth of multimedia information, and are valuable as archived material. In order to extract speech scenes from news videos, there is an approach that uses the position and size of a face region. However, it is difficult to extract them with only such approach, since news videos contain non-speech scenes where the speaker is not the subject, such as narrated scenes. To solve this problem, we propose a method to discriminate between speech scenes and narrated scenes based on the co-occurrence between a subject's lip motion and the speaker's voice. The proposed method uses lip shape and degree of lip opening as visual features representing a subject's lip motion, and uses voice volume and phoneme as audio feature representing a speaker's voice. Then, the proposed method discriminates between speech scenes and narrated scenes based on the correlations of these features. We report the results of experiments on videos captured in a laboratory condition and also on actual broadcast news videos. Their results showed the effectiveness of our method and the feasibility of our research goal.
Shogo Kumagai, Keisuke Doman, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ISM6
2011 Intelligent traffic sign detector: Adaptive learning based on online gathering of training samples
abstract
This paper proposes an intelligent traffic sign detector using adaptive learning based on online gathering of training samples from in-vehicle camera image sequences. To detect traffic signs accurately from in-vehicle camera images, various training samples of traffic signs are needed. In addition, to reduce false alarms, various background images should also be prepared before constructing the detector. However, since their appearances vary widely, it is difficult to obtain them exhaustively by manual intervention. Therefore, the proposed method simultaneously obtains both traffic sign images and background images from in-vehicle camera images. Especially, to reduce false alarms, the proposed method gathers background images that were easily mis-detected by a previously constructed traffic sign detector, and re-trains the detector by using them as negative samples. By using retrospectively tracked traffic sign images and background images as positive and negative training samples, respectively, the proposed method constructs a highly accurate traffic sign detector automatically. Experimental results showed the effectiveness of the proposed method.
Daisuke Deguchi, Daisuke Shirasuna, Keisuke Doman, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium5
2011 Estimation of traffic sign visibility considering temporal environmental changes for smart driver assistance
abstract
We propose a visibility estimation method for traffic signs considering temporal environmental changes, as a part of work for the realization of nuisance-free driver assistance systems. Recently, the number of driver assistance systems in a vehicle is increasing. Accordingly, it is becoming important to sort out appropriate information provided from them, because providing too much information may cause driver distraction. To solve such a problem, we focus on a visibility estimation method for controlling the information according to the visibility of a traffic sign. The proposed method sequentially captures a traffic sign by an in-vehicle camera, and estimates its accumulative visibility by integrating a series of instantaneous visibility. By this way, even if the environmental conditions may change temporally and complicatedly, we can still accurately estimate the visibility that the driver perceives in an actual traffic scene. We also investigate the performance of the proposed method and show its effectiveness.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
Intelligent Vehicles Symposium6
2011 Road image update using in-vehicle camera images and aerial image
abstract
Road image is becoming important for several applications such as car navigation systems, traffic environment research, city modeling. Usually, a road image can be obtained from an aerial image but the resolution of the aerial image is often low, or it contains occlusions by obstacles. Therefore, the update of road image is required. In this paper, we propose a road image mosaicing method using in-vehicle camera images and an aerial image. We first perform image registration of road regions between these images, and then, we generate a large road image by performing image mosaicing of road regions in invehicle camera images. In an experiment, we achieved resolution improvement and occlusions removal, and also succeeded in update of a large road image.
Masafumi Noda, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase, Yoshiko Kojima, Takashi Naito
Intelligent Vehicles Symposium5
2011 3-D line segment reconstruction using an in-vehicle camera for free space detection
abstract
Free space detection is very important for vehicle navigation and safe driving. 3-D line segment reconstruction of a street is important for the free space detection because a street-view includes many line segments. For the free space detection, we propose a method for reconstructing 3-D line segments in a streetscape using a monocular in-vehicle camera. The 3-D reconstruction of the line segments is achieved by using each three images from an image sequence. Once accurate camera poses of these images are obtained, one of the remaining crucial problems is to match the line segments between the images correctly. A strategy for finding correspondence of the line segments is as follows: First, the correspondences of line segment candidates are searched by using a two-view constraint. However, the two-view constraint has difficulty on determining an unique correspondence geometrically. Therefore, the candidates of the line segment correspondences are reduced using a three-view constraint. In order to improve the accuracy, the proposed method exploits a color feature of the line segment and a preliminary knowledge of the vehicle motion. Finally, the line segments are reconstructed using the correspondences. From an experimental result, we confirmed the effectiveness of the proposed method. Application to the free space detection demonstrated the usefulness of the reconstructed line segments.
Hiroyuki Uchiyama, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
Intelligent Vehicles Symposium5
2011 Scene segmentation of wedding party videos by scenario-based matching with example videos
abstract
We propose a method for scene segmentation of a wedding party video. Recently, it has become popular to take videos of a wedding ceremony and its party. Especially, because of its length, each scene of a wedding party video needs to be indexed with each event for efficient browsing. The proposed method segments a wedding party video into scenes of events by scenario- based matching with example videos that are synthesized by combining scenes from other wedding party videos according to a scenario.
Kazuki Sawai, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ACM Multimedia5
2011 Video CooKing: Towards the Synthesis of Multimedia Cooking Recipes
Keisuke Doman, Cheng Ying Kuai, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
MMM (2)5
2010 Classification of Near-Duplicate Video Segments Based on Their Appearance Patterns
abstract
We propose a method that analyzes the structure of a large volume of general broadcast video data by the appearance patterns of near-duplicate video segments. We define six classification rules based on the appearance patterns of near-duplicate video segments according to their roles, and evaluated them over more than 1,000 hours of actual broadcast video data.
Ichiro Ide, Yuji Shamoto, Daisuke Deguchi, Tomokazu Takahashi, Hiroshi Murase
ICPR5
2010 Efficient Facial Attribute Recognition with a Spatial Codebook
abstract
There is a large number of possible facial attributes such as hairstyle, with/without glasses, with/without mustache, etc. Considering large number of facial attributes and their combinations, it is difficult to build attributes classifiers for all possible combinations needed in various applications, especially at the designing stage. To tackle this important and challenging problem, we propose a novel efficient facial attributes recognition algorithm using a learned spatial codebook. The Maximum Entropy and Maximum Orthogonality (MEMO) criterion is followed to learn the spatial codebook. With a spatial codebook constructed at the designing stage, attribute classifiers can be trained on demand with a small number of exemplars with high accuracy on the testing data. Meanwhile, up to 600 times speedup is achieved in the on-demand training process, compared to current state-of-the-art method. The effectiveness of the proposed method is supported by convincing experimental results.
Yoshihisa Ijiri, Shihong Lao, Tony X. Han, Hiroshi Murase
ICPR4
2010 Region-Based Image Transform for Transition Between Object Appearances
abstract
We propose a method of region-based image transform to achieve accurate transition between object appearances. A view-transition model (VTM) is one of the statistical methods that learn appearance transition from a sample image dataset of a large number of objects with various appearances. However, the VTM method has a practical problem that the appearance transition cannot be performed accurately if a sufficient number of learning samples is not available in the dataset. To cope with the problem, the proposed method first determines the regions of input and output images whose pixel values mutually affect each other during appearance transition, then transforms iteratively between partial images in the regions. We conducted experiments using actual image datasets. The results show that the proposed method could accurately transform appearances compared with the VTM method.
Tomokazu Takahashi, Yuki Kono, Ichiro Ide, Hiroshi Murase
ICPR4
2010 Removal of Moving Objects from a Street-View Image by Fusing Multiple Image Sequences
abstract
We propose a method to remove moving objects from an in-vehicle camera image sequence by fusing multiple image sequences. Driver assistance systems and services such as Google Street View require images containing no moving object. The proposed scheme consists of three parts: (i) collection of many image sequences along the same route by using vehicles equipped with an omni-directional camera, (ii) temporal and spatial registration of image sequences, and (iii) mosaicing partial images containing no moving object. Experimental results show that 97.3% of the moving object area could be removed by the proposed method.
Hiroyuki Uchiyama, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICPR5
2010 Multimedia Supplementation to a Cooking Recipe Text for Facilitating Its Understanding to Inexperienced Users
abstract
Assisting culinary activities for inexperienced users has been considered as an important task in most existing works in the field. On the other hand, recipe texts are becoming available on the Internet in increasing numbers. However, they tend to be written simply by mostly non-professional people, and thus are sometimes difficult for an inexperienced person to follow the steps and manage to cook as they are supposed to. In this paper, we propose a method that detects difficult descriptions for an inexperienced user in an existing text recipe, and supplements them with multimedia contents including text information extracted from a large number of recipes, and also images and video clips on certain kinds of cooking operations, to facilitate the understanding of the recipe. Experimental results showed promising ability of the proposed method to assist inexperienced users understand the descriptions in a recipe.
Ichiro Ide, Yuka Shidochi, Yuichi Nakamura 0001, Daisuke Deguchi, Tomokazu Takahashi, Hiroshi Murase
ISM6
2010 Estimation of traffic sign visibility toward smart driver assistance
abstract
We propose a visibility estimation method for traffic signs as part of work for realization of nuisance-free driving safety support systems. Recently, the number of driving safety support systems in a car has been increasing. As a result, it is becoming important to select appropriate information from them for safe and comfortable driving because too much information may cause driver distraction and may increase the risk of a traffic accident. One of the approaches to avoid such a problem is to alert the driver only with information which could easily be missed. Therefore, to realize such a system, we focus on estimating the visibility of traffic signs. The proposed method is a model-based method that estimates the visibility of traffic signs focusing on the difference of image features between a traffic sign and its surrounding region. In this paper, we investigate the performance of the proposed method and show its effectiveness.
Keisuke Doman, Daisuke Deguchi, Tomokazu Takahashi, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yukimasa Tamatsu
Intelligent Vehicles Symposium6
2010 A Hilbert warping method for handwriting gesture recognition
Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
Pattern Recognit.4
2009 Low-Resolution Character Recognition by Video-Based Super-Resolution
abstract
In this paper, we propose a method for recognizing low-resolution characters using a super-resolution technique. Although portable digital cameras can be used for camera based character recognition, the captured images contain several types of noises which make the recognition task difficult. We introduce a phase of super-resolution before the recognition to enhance the resolution of images obtained from a video. The proposed method uses the subspace method for the recognition of characters which are integrated from multiple low-resolution characters by the super-resolution technique. Experimental results show that the proposed method improves the recognition accuracy; we confirmed that the recognition rate for the input size of 7 times 7 pixels was 90.35%, and for the input size of 9 times 9 pixels was 99.97%.
Ataru Ohkura, Daisuke Deguchi, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICDAR5
2009 Adaptive division of feature space for rapid detection of near-duplicate video segments
abstract
Near-duplicate video detection is becoming a core-technology for analyzing the structure of a large-scale video archive. It, however, is naturally an O(n2) problem, where n is a value proportional to the total length of an input video stream. We have previously challenged this time-consuming task by reducing the cost required for each of the O(n2) comparisons. This paper, on the other hand, proposes a method that reduces the number of comparisons by adaptively dividing the feature space according to the distribution of feature points.
Ichiro Ide, Shugo Suzuki, Tomokazu Takahashi, Hiroshi Murase
ICME4
2009 Labeling News Topic Threads with Wikipedia Entries
abstract
Wikipedia is a famous online encyclopedia. However most Wikipedia entries are mainly explained by text, so it will be very informative to enhance the contents with multimedia information such as videos. Thus we are working on a method to extend information of Wikipedia entries by means of broadcast videos which explain the entries. In this work, we focus especially on news videos and Wikipedia entries about news events. In order to extend information of Wikipedia entries, it is necessary to link news videos and Wikipedia entries. So the main issue will be on a method that labels news videos with Wikipedia entries automatically. In this way, explanations could be more detailed with news videos can be exhibited, and the context of the news events should become easier to understand. Through experiments, news videos were accurately labeled with Wikipedia entries with a precision of 86% and a recall of 79%.
Tomoki Okuoka, Tomokazu Takahashi, Daisuke Deguchi, Ichiro Ide, Hiroshi Murase
ISM5
2009 A Multimodal Constellation Model for Object Category Recognition
Yasunori Kamiya, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
MMM4
2008 A Hilbert Warping Algorithm for Recognizing Characters from Moving Camera
abstract
We present a method for recognizing characters from image sequences captured by moving camera. In the proposed method, the sequence of the captured images is compared with those of reference character patterns using the concept of analytic signal. Since the captured image sequence can be nonlinearly warped along the time axis due to the movement of a hand-held camera, phase synchronization of two analytic signals is used for the alignment of two image sequences. Hilbert transform is used to convert all the image sequences into analytic signals whose phases are supposed to be increasing. Experimental results showed the usefulness of the proposed phase-based alignment algorithm.
Hiroyuki Ishida, Ichiro Ide, Hiroshi Murase, Tomokazu Takahashi
Document Analysis Systems3
2008 A Hilbert warping method for camera-based finger-writing recognition
abstract
We propose a time-warping algorithm for recognizing finger actions by a camera. In the proposed method, an input image sequence is aligned to the reference sequences by phase-synchronization of the analytic signals, and then classified by comparing the cumulative distances. A major benefit of this method is that over-fitting to sequences of incorrect categories is restricted. The proposed method exhibited high recognition accuracy in finger-writing character recognition.
Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
ICPR4
2008 Eigenspace interpolation for appearance-based object recognition
abstract
An eigenspace interpolation method smoothly interpolates between two different eigenspaces using high dimensional rotation. However, up to now its effectiveness in object recognition and the validity of the interpolation algorithm have not been discussed sufficiently. We therefore propose an appearance-based object recognition method combining the eigenspace interpolation method and a subspace method. We conducted face recognition experiments using images captured from multiple camera positions with various illumination conditions. Experimental results demonstrate the effectiveness of the proposed method and the validity of the interpolation algorithm.
Tomokazu Takahashi, Lina, Ichiro Ide, Yoshito Mekada, Hiroshi Murase
ICPR5
2008 Cross-Lingual Retrieval of Identical News Events by Near-Duplicate Video Segment Detection
Akira Ogawa, Tomokazu Takahashi, Ichiro Ide, Hiroshi Murase
MMM4
2008 Recognition of camera-captured low-quality characters using motion blur information
Hiroyuki Ishida, Tomokazu Takahashi, Ichiro Ide, Yoshito Mekada, Hiroshi Murase
Pattern Recognit.5
2008 A Quick Search Method for Audio Signals Based on a Piecewise Linear Representation of Feature Trajectories
abstract
This paper presents a new method for a quick similarity-based search through long unlabeled audio streams to detect and locate audio clips provided by users. The method involves feature-dimension reduction based on a piecewise linear representation of a sequential feature trajectory extracted from a long audio stream. Two techniques enable us to obtain a piecewise linear representation: the dynamic segmentation of feature trajectories and the segment-based Karhunen-L\'{o}eve (KL) transform. The proposed search method guarantees the same search results as the search method without the proposed feature-dimension reduction method in principle. Experiment results indicate significant improvements in search speed. For example the proposed method reduced the total search time to approximately 1/12 that of previous methods and detected queries in approximately 0.3 seconds from a 200-hour audio database.
Akihiro Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
IEEE Trans. Speech Audio Process.4
2007 Interpolation Between Eigenspaces Using Rotation in Multiple Dimensions
Tomokazu Takahashi, Lina, Ichiro Ide, Yoshito Mekada, Hiroshi Murase
ACCV (2)5
2007 Genre-Adaptive Near-Duplicate Video Segment Detection
abstract
This paper proposes a fast and accurate method to detect all near-duplicate segments in a video stream. To reduce the computation time while ensuring the detection accuracy equivalent to that by brute-force frame-by-frame comparison, a two-step detection method is proposed; a fast but rough detection applied in a compressed feature vector space spanned by the result of a PC A, followed by confirmation of candidates in the original high dimension space. The results show that the proposed method accelerates the detection by more than 1,000 times while maintaining the detection accuracy. We also propose an entropy-based pixel selection scheme to generate feature vectors optimized for comparison of video segments within programs with mostly common pictures. The results show that the proposed scheme eliminates the false positives drastically, which should lead to even faster detection.
Ichiro Ide, Kazuhiro Noda, Tomokazu Takahashi, Hiroshi Murase
ICME4
2007 mediaWalker: a video archive explorer based on time-series semantic structure
abstract
We introduce a video browsing interface 'mediaWalker' that lets users explore a news video archive based on a time-series semantic structure; the 'topic thread' structure. The interface lets users efficiently track up and down the development of news in an archive with more than 1,000 hours of video.
Ichiro Ide, Tomoyoshi Kinoshita, Tomokazu Takahashi, Shin'ichi Satoh 0001, Hiroshi Murase
ACM Multimedia5
2006 Spatiotemporal Density Feature Analysis to Detect Liver Cancer from Abdominal CT Angiography
Yoshito Mekada, Yuki Wakida, Yuichiro Hayashi, Ichiro Ide, Hiroshi Murase
ACCV (2)5
2006 Conversation Scene Analysis with Dynamic Bayesian Network Basedon Visual Head Tracking
abstract
A novel method based on a probabilistic model for conversation scene analysis is proposed that can infer conversation structure from video sequences of face-to-face communication. Conversation structure represents the type of conversation such as monologue or dialogue, and can indicate who is talking/listening to whom. This study assumes that the gaze directions of participants provide cues for discerning the conversation structure, and can be identified from head directions. For measuring head directions, the proposed method newly employs a visual head tracker based on sparse-template condensation. The conversation model is built on a dynamic Bayesian network and is used to estimate the conversation structure and gaze directions from observed head directions and utterances. Visual tracking is conventionally thought to be less reliable than contact sensors, but experiments confirm that the proposed method achieves almost comparable performance in estimating gaze directions and conversation structure to a conventional sensor-based method
Kazuhiro Otsuka, Junji Yamato, Yoshinao Takemae, Hiroshi Murase
ICME4
2005 Automated Nomenclature of Bronchial Branches Extracted from CT Images and Its Application to Biopsy Path Planning in Virtual Bronchoscopy
Kensaku Mori, Sinya Ema, Takayuki Kitasaka, Yoshito Mekada, Ichiro Ide, Hiroshi Murase, Yasuhito Suenaga, Hirotsugu Takabatake, Masaki Mori, Hiroshi Natori
MICCAI (2)6
2003 Dynamic-segmentation-based feature dimension reduction for quick audio/video searching
abstract
We propose a new feature dimension reduction method for multimedia search. The main technique in the method is dynamic segmentation that partitions sequential feature trajectories dynamically. While dynamic segmentation reduces the average dimensionality and accelerates the search, it requires huge amount of calculation. Thus, our method quickly executes suboptimal partitioning of the trajectories by using the discreteness of dimension changes. This guarantees the optimal amount of calculation to derive the suboptimal partitioning under the condition that the dimension monotonously increases as the segment length increases. The experiment shows that our method is over 10 times faster than a straightforward dynamic segmentation method.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP (3)4
2003 A fast search algorithm for background music signals based on the search for numerous small signal components
abstract
The paper proposes a method for detecting and locating a known music signal in a long audio stream. Unlike existing methods, ours assumes that the music is used as background music (BGM) and overlapped by another sound such as speech and that the interfering sound is typically louder than the target music. The proposed method is based on time-series active search, which is a quick signal search method reported earlier (Kashino, K. et al., Proc. ICASSP-99, vol.VI, 1999). To realize the BGM search, however, a novel extension is introduced. That is, the music signal is first decomposed into a number of small time-frequency regions, and the search is carried out for each of those components. The results of the search are then integrated based on a voting scheme to find the target music locations. Experiments show that an accurate search is possible when SNR is -5 dB and that the search completes in about 8 s for a 30 min stored signal.
Hidehisa Nagano, Kunio Kashino, Hiroshi Murase
ICASSP (5)3
2003 Small cylindrical display for anthropomorphic agents
abstract
This paper describes a small cylindrical display for an anthropomorphic agent that communicates with users in a three-dimensional (3D) environment. Conventional displays for anthropomorphic agents are designed for a single user or a fixed viewing direction. We suggest a new small cylindrical display that is visible from the outside in any direction. Image data for the projector is calculated by the image-warping method. Light emitted from the image projector is reflected by a spherical mirror and projected on the inside of the rear of the cylindrical screen. This cylindrical display enables 3D agent actions, such as head turning. We evaluated our display by measuring pixel brightness in different situations. We demonstrate an application of anthropomorphic agents and a telecommunications system using omnidirectional images.
Takahito Kawanishi, Masaru Tsuchida, Shigeru Takagi, Hiroshi Murase
ICME4
2003 Dynamic-segmentation-based feature dimension reduction for quick audio/video searching
abstract
We propose a new feature dimension reduction method for multimedia search. The main technique in the method is dynamic segmentation that partitions sequential feature trajectories dynamically. While dynamic segmentation reduces the average dimensionality and accelerates the search, it requires huge amount of calculation. Thus, our method quickly executes suboptimal partitioning of the trajectories by using the discreteness of dimension changes. This guarantees the optimal amount of calculation to derive the suboptimal partitioning under the condition that the dimension monotonously increases as the segment length increases. The experiment shows that our method is over 10 times faster than a straightforward dynamic segmentation method.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICME4
2003 A fast search algorithm for background music signals based on the search for numerous small signal components
abstract
This paper proposes a method for detecting and locating a known music signal in a long audio stream. Unlike existing methods, ours assumes that the music is used as background music (BGM) and overlapped by another sound such as speech and that the interfering sound is typically louder than the target music. The proposed method is based on time-series active search, which is a quick signal search method reported earlier. To realize the BGM search, however, a novel extension is introduced. That is, the music signal is firstly decomposed into a number of small time-frequency regions, and the search is carried out for each of those components. The results of the search are then integrated based on a voting scheme to find the target music locations. Experiments show that accurate search is possible when SNR is -5 dB and that the search completes in about 8 s for a 30-m stored signal.
Hidehisa Nagano, Kunio Kashino, Hiroshi Murase
ICME3
2003 Unsupervised recognition of multi-view face sequences based on pairwise clustering with attraction and repulsion
Bisser Raytchev, Hiroshi Murase
Comput. Vis. Image Underst.2
2003 Unsupervised face recognition by associative chaining
Bisser Raytchev, Hiroshi Murase
Pattern Recognit.2
2003 A quick search method for audio and video signals based on histogram pruning
abstract
This paper proposes a quick method of similarity-based signal searching to detect and locate a specific audio or video signal given as a query in a stored long audio or video signal. With existing techniques, similarity-based searching may become impractical in terms of computing time in the case of searching through long-running (several-days' worth of) signals. The proposed algorithm, which is referred to as time-series active search, offers significantly faster search with sufficient accuracy. The key to the acceleration is an effective pruning algorithm introduced in the histogram matching stage. Through the pruning, the actual number of matching calculations can be reduced by 200 to 500 times compared with exhaustive search while guaranteeing exactly the same search result. Experiments show that the proposed method can correctly detect and locate a 15-s signal in a 48-h recording of TV broadcasts within 1 s, once the feature vectors are calculated and quantized. As extentions of the basic algorithm, efficient AND/OR search methods for searching for multiple query signals and a feature dithering method for coping with signal distortion are also discussed.
Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
IEEE Trans. Multim.3
2002 A quick search method for multimedia signals using feature compression based on piecewise linear maps
abstract
We propose a quick algorithm for multimedia signal search. The algorithm comprises two techniques: feature compression based on piecewise linear maps and distance bounding to efficiently limit the search space. When compared with existing multimedia search techniques, they greatly reduce the computational cost required in searching. Although feature compression is employed in our method, our bounding technique mathematically guarantees the same recall rate as the search based on the original features; no segment to be detected is missed. Experiments indicate that the proposed algorithm is approximately 10 times faster than and as accurate as an existing fast method maitaining the same search accuracy.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP4
2002 VQ-faces - unsupervised face recognition from image sequences
abstract
We propose a new method for unsupervised face recognition - VQ-faces, which operates on a sequential stream of face images and is able to handle both frontal and side-view faces at the same time. The method consists of two parts: in the first part, the VQ-faces are calculated as prototype vectors of local areas in image-space, coding for different face-views (i.e. a "view codebook" is generated), while in the second part the fact that each face-sequence corresponds to a single person (temporal constraint) is used to cluster the different sequences into face categories, using a combinatorial optimization process which maximizes an objective function, reducing the global representation error (distortion) by the VQ-faces as sequences are grouped together. The method was tested on real-world data gathered over a period of several months and including both frontal and side-view faces from 17 different subjects, achieving correct self-organization rate of 85.4%. The applicability of the proposed method is not limited to face recognition - it can be easily applied to other problems involving multi-view object recognition.
Bisser Raytchev, Hiroshi Murase
ICIP (2)2
2002 Fast music retrieval using polyphonic binary feature vectors
abstract
We propose a method for retrieving similar music from a polyphonic-music audio database using a polyphonic audio signal as a query. In this task, we must consider similarities among polyphonic signals of the music, and achieve quick retrieval. Therefore, we first introduce the polyphonic binary feature vector to represent the presence of multiple notes. This feature is suitable for the search based on the similarities among polyphonic audio signals. Then, we propose a new search method, which is quicker than the exhaustive use of DP matching. The search is accelerated using a "similarity matrix" to limit the search space. Experiments using a test database containing 216 music pieces show that the search accuracy of the proposed feature is 89%, which is approximately 26% higher than that of the conventional spectrum feature. It is also shown that the new search method retrieves similar music without significant accuracy degradation as well as the exhaustive search does and the computational complexity of the new search method is about 1/4 that of exhaustive search.
Hidehisa Nagano, Kunio Kashino, Hiroshi Murase
ICME (1)3
2001 Unsupervised Face Recognition from Image Sequences Based on Clustering with Attraction and Repulsion
abstract
We propose a new method for unsupervised face recognition from time-varying sequences of face images obtained in real-world environments. Two types of forces, attraction and repulsion, operate across the spatio-temporal facial manifolds, to autonomously organize the data without relying on any category-specific information provided in advance. Experiments with real-world data gathered over a period of several months and including both frontal and side-view faces were used to evaluate the method and encouraging results were obtained The proposed method can be used in video surveillance systems or for content-based information retrieval.
Bisser Raytchev, Hiroshi Murase
CVPR (2)2
2001 Very quick audio searching: introducing global pruning to the Time-Series Active Search
abstract
Previously, we proposed a histogram-based quick signal search method called Time-Series Active Search (TAS). TAS is a method of searching through long audio or video recordings for a specified segment, based on signal similarity. TAS is fast; it can search through a 24-hour recording in 1 second after a query-independent preprocessing. However, an even faster method is required when we consider a huge amount of audio archives, for example a month's worth of recordings. Thus, we propose a preprocessing method that significantly accelerates TAS. The core part of this method comprises a global histogram clustering of long signals and a pruning scheme using those clusters. Tests using broadcast recording indicate that the proposed algorithm achieves a search speed approximately 3 to 30 times faster than TAS. In these tests, the search results are exactly the same as with TAS.
Akisato Kimura, Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICASSP4
2001 Robust Feature Extraction Based on Run-Length Compensation for Degraded Handwritten Character Recognition
abstract
Conventional features are robust for recognizing either deformed or degraded characters. This paper proposes a feature extraction method that is robust for both of them. Run-length compensation is introduced for extracting approximate directional run-lengths of strokes from degraded handwritten characters. This technique is applied to the conventional feature vector based on directional run-lengths. Experiments for handwritten characters with additive or subtractive noise show that the proposed feature is superior to conventional ones over a wide range of the degree of noise.
Minoru Mori, Minako Sawaki, Norihiro Hagita, Hiroshi Murase, Naoki Mukawa
ICDAR4
2001 A method for robust and quick video searching using probabilistic dither-voting
abstract
We propose a quick and accurate search method for detecting a query signal from long video recordings The method is based on the time-series active search, which is a quick searching method for audio and video signals that we previously proposed. Time-series active search is based on a histogram matching scheme and an efficient pruning mechanism, and therefore, it was very quick. We found, however, that the accuracy sometimes deteriorates when it is applied to searches through long video archives that are composed of many similar video images or those containing feature distortions caused by video dubbing or low-bit-rate compression. The problem arises from (1) insufficient capability of representing features and (2) feature distortions. Thus, the method proposed here uses LBG-based VQ to improve the capacity to represent features and probabilistic dither-voting to improve robustness with respect to feature distortions. The experiments prove the effects of the proposed method.
Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICIP (2)3
2001 Dynamic Active Search for quick object detection with pan-tilt-zoom camera
abstract
This paper proposes a search method for detecting known objects quickly in 3D environments with a pan-tilt-zoom camera. In our previous work, we proposed an algorithm named Active Search that greatly reduces the number of calculations required to obtain a match between a reference object and an input image using color histograms. We describe two improvements we have made to Active Search for such practical applications as robots and surveillance. First, we increased the robustness as regards the color histogram changes that result from different lighting conditions and camera angles by using multiple reference images and a pixel color vector quantization. Second, we reduced the number of camera operations (pan, tilt and zoom) by using a best-direction-first and upper bound pruning strategies. We call this camera control Dynamic Active Search. Experiments show an improvement in object detection accuracy and a 78% reduction in detection time.
Takahito Kawanishi, Hiroshi Murase, Shigeru Takagi
ICIP (3)2
2001 Unsupervised face recognition from image sequences
abstract
We propose a novel method for unsupervised face recognition from time-varying sequences of face images obtained in real-world environments. The method utilizes the higher level of sensory variation contained in the input image sequences to autonomously organize the data in an incrementally built graph structure, without relying on category-specific information provided in advance. This is achieved by "chaining" together similar views across the spatio-temporal facial manifolds by two types of connecting edges depending on a local measure of similarity. Experiments with real-world data gathered over a period of several months and including both frontal and side-view faces from 17 different subjects were used to test the method, achieving a correct self-organization rate of 88.6%. The proposed method can be used in video surveillance systems or for content-based information retrieval.
Bisser Raytchev, Hiroshi Murase
ICIP (1)2
2001 Expressing Personality of Interface Agents by Gaze
Atsushi Fukayama, Minako Sawaki, Takehiko Ohno, Hiroshi Murase, Norihiro Hagita, Naoki Mukawa
INTERACT4
2001 Unsupervised Learning of Faces for Human-Computer Interfaces
Bisser Raytchev, Hiroshi Murase
INTERACT2
2000 Feature Fluctuation Absorption for a Quick Audio Retrieval from Long Recordings
abstract
Kashino et al. proposed (1999) a histogram-based quick signal search method called time-series active search (TAS). TAS has only been effective in the exact matching case, where the segments to be detected are assumed to be exactly same as the reference signal. Here, we extend the method so that it is applicable even if the features fluctuate. In addition to the feature modification, feature dithering is discussed to absorb feature fluctuations. Efficient time-scaled search is also investigated to cope with variations of the reference signal duration. Tests using broadcast recordings show that the extended method improves the accuracy in nonexact-matching tasks such as hand-clap detection and word spotting in a single-speaker's narration. The tests also show the speed-ups by pruning introduced in the time-scaled search.
Kunio Kashino, Takayuki Kurozumi, Hiroshi Murase
ICPR3
2000 Automatic Acquisition of Context-based Image Templates for Degraded Character Recognition in Scene Images
abstract
Proposes a method for adaptively acquiring templates for degraded characters in scene images. Characters in scene images are often degraded because of poor printing and viewing conditions. To cope with the degradation problem, we proposed the idea of "context-based image templates" which include neighboring characters of parts thereof and so represent more contextual information than single-letter templates. However, our previous method manually selects the learning samples to make the context-based image templates and is time-consuming. Therefore, we attempt to make the context-based image templates automatically from single-letter templates and learning text-line images. The context-based image templates are iteratively created using the k-nearest neighbor rule. Experiments with 3,467 alpha-numeric characters in nine bookshelf images show that the high recognition rates for test samples possible with this method asymptotically approach those achieved with manual selection.
Minako Sawaki, Hiroshi Murase, Norihiro Hagita
ICPR2
1999 Time-series active search for quick retrieval of audio and video
abstract
This paper proposes a search method that can quickly detect and locate known sound (video) in a long audio (video) stream. The method is based on active search. Active search reduces the number of candidate matches between reference and input signals by approximately 10 to 100 times compared to exhaustive search, while guaranteeing the same retrieval accuracy. We proposed a quick search method in Smith et al. (1998), and here we focus on improvement of the accuracy. Thus the feature used has been extended to the audio power spectrum and temporal division of the histogram windows has been introduced to incorporate time information. Tests carried out under practical circumstances clearly show the accuracy improvement. The proposed method is still so fast that it can correctly retrieve a 15-s commercial in a 6-h recording of TV broadcasting within 2 s, once the features are calculated.
Kunio Kashino, Gavin Smith, Hiroshi Murase
ICASSP3
1999 Multi-category classification by kernel based nonlinear subspace method
abstract
The kernel based nonlinear subspace (KNS) method is proposed for multi-class pattern classification. This method consists of the nonlinear transformation of feature spaces defined by kernel functions and subspace method in transformed high-dimensional spaces. The support vector machine, a nonlinear classifier based on a kernel function technique, shows excellent classification performance, however, its computational cost increases exponentially with the number of patterns and classes. The linear subspace method is a technique for multi-category classification, but it fails when the pattern distribution has nonlinear characteristics or the feature space dimension is low compared to the number of classes. The proposed method combines the advantages of both techniques and realizes multi-class nonlinear classifiers with better performance in less computational time. We show that a nonlinear subspace method can be formulated by nonlinear transformations defined through kernel functions and that its performance is better than that obtained by conventional methods.
Eisaku Maeda, Hiroshi Murase
ICASSP2
1999 Character Recognition in Bookshelf Images using Context-based Image Templates
abstract
This paper proposes a method for recognizing degraded characters in bookshelf images captured by a digital camera. We adopt displacement matching and templates that include neighboring characters or parts thereof to cope with the degradation. The templates are referred as context-based image templates, since they offer more contextual information than single-letter templates. Such templates are effective wherever there is a restricted word set, such as journal titles, year, month, volume, and number. Experiments with 3,468 characters in nine bookshelf images show that this method achieves a higher recognition rare (96.3%) than single-letter templates (88.4%).
Minako Sawaki, Hiroshi Murase, Norihiro Hagita
ICDAR2
1999 A sound source identification system for ensemble music based on template adaptation and music stream extraction
Kunio Kashino, Hiroshi Murase
Speech Commun.2
1998 Active Viewpoint Control for Shape from Occluding Contours
Takashi Akutsu, Kenichi Arakawa, Hiroshi Murase
ACCV (1)3
1998 Music recognition using note transition context
abstract
As a typical example of sound-mixture recognition, the recognition of ensemble music is addressed. Here music recognition is defined as recognizing the pitch and the name of an instrument for each musical note in monaural or stereo recordings of real music performances. The first key part of the proposed method is adaptive template matching that can cope with variability in musical sounds. This is employed in the hypothesis-generation stage. The second key part of the proposed method is musical context integration based on the probabilistic networks. This is employed in the hypothesis-verification stage. The evaluation results clearly show the advantages of these two processes.
Kunio Kashino, Hiroshi Murase
ICASSP2
1998 Quick audio retrieval using active search
abstract
This paper discusses a method to search quickly through broadcast audio data to detect and locate known sounds using reference templates, based on the active search algorithm and histogram modeling of zero-crossing features. Active search reduces the number of candidate matches between reference and test template by up to 36 times compared to exhaustive search, while still remaining optimal. Computation is further reduced by using computationally inexpensive zero-crossing features. The method is robust against white noise addition down to 20 dB signal-to-noise ratios and digitization noise.
Gavin Smith, Hiroshi Murase, Kunio Kashino
ICASSP2
1998 Robust Object Extraction with Illumination-Insensitive Color Descriptions
Chie Hashizume, Vasudevan V. Vinod, Hiroshi Murase
ICIP (3)3
1998 Character recognition in bookshelf images by automatic template selection
abstract
Presents a multiple-dictionary method for recognizing low-quality characters in scene images. First, the environmental conditions of an input image are estimated using an initial dictionary. Then, a relevant dictionary from multiple dictionaries reflecting different environmental conditions is automatically selected from the estimation and used for recognition. Experiments are made for characters in images of bookshelves. The results show that the proposed method achieves a higher recognition rate (89.8%) than that obtained by using a single dictionary (76.4%). Furthermore, recognition accuracy improves from 89.8% to 95.2% using contextual postprocessing.
Minako Sawaki, Hiroshi Murase, Norihiro Hagita
ICPR2
1998 Parametric Feature Detection
Simon Baker, Shree K. Nayar, Hiroshi Murase
Int. J. Comput. Vis.3
1998 Coarse-to-fine adaptive masks for appearance matching of occluded scenes
Jeff L. Edwards, Hiroshi Murase
Mach. Vis. Appl.2
1998 Image retrieval using efficient local-area matching
Vasudevan V. Vinod, Hiroshi Murase
Mach. Vis. Appl.2
1997 Appearance Matching of Occluded Objects Using Coarse-to-fine Adaptive Masks
abstract
In this paper, we discuss an appearance matching technique for the interpretation of color scenes containing occluded objects. Dealing with occlusions is very difficult, and we have explored the use of an iterative, coarse-to-fine correlation-based method that uses hypothesized occlusion events to modify the scene-to-template similarity measure at run-time. Specifically, a binary mask is used to adaptively exclude regions of the template image from the correlation computation. At each iteration, these masks are adjusted based on higher resolution scene data and the occluding interactions between multiple object hypotheses. We present results which demonstrate the technique is reasonably robust over a large database of color test scenes containing objects at a variety of scales, and tolerates minor object rotations and global illumination variations.
Jeff L. Edwards, Hiroshi Murase
CVPR2
1997 A Music Stream Segregation System Based on Adaptive Multi-Agents
Kunio Kashino, Hiroshi Murase
IJCAI2
1997 Focused color intersection with efficient searching for object extraction
abstract
We propose focused color intersection with efficient searching for identifying and extracting the objects in a complex scene based on color similarity. The method matches the models against different parts of a scene, called focus regions, using normalized color histogram intersection. The best matching focus region is determined by an efficient search strategy employing upper bound pruning. This search strategy, called active search, concentrates its effort on parts of the scene having high similarity with the object. Consequently, it achieves a large reduction in computational effort without sacrificing accuracy. An efficient algorithm for evaluating the color histogram intersection between a model and a focus region is also given. Experiments conducted demonstrate that multiple known objects in complex scenes can be extracted by this process. The method is stable against scale changes, two-dimensional rotation, moderate changes in shape and partial occlusion.
Vasudevan V. Vinod, Hiroshi Murase
Pattern Recognit.2
1997 Detection of 3D objects in cluttered scenes using hierarchical eigenspace
abstract
This paper proposes a novel method to detect three-dimensional objects in arbitrary poses and sizes from a complex image and to simultaneously measure their poses and sizes using appearance matching. In the learning stage, for a sample object to be learned, a set of images is obtained by varying pose and size. This large image set is compactly represented by a manifold in compressed subspace spanned by eigenvectors of the image set. This representation is called the parametric eigenspace representation. In the object detection stage, a partial region in an input image is projected to the eigenspace, and the location of the projection relative to the manifold determines whether this region belongs to the object, and what its pose is in the scene. This process is sequentially applied to the entire image at different resolutions. Experimental results show that this method accurately detects the target objects.
Hiroshi Murase, Shree K. Nayar
Pattern Recognit. Lett.1
1996 Parametric Feature Detection
abstract
We propose an algorithm to automatically construct feature detectors for arbitrary parametric features. To obtain a high level of robustness we advocate the use of realistic multi-parameter feature models and incorporate optical and sensing effects. Each feature is represented as a densely sampled parametric manifold in a low dimensional subspace of a Hilbert space. During detection, the brightness distribution around each image pixel is projected into the subspace. If the projection lies sufficiently close to the feature manifold, the feature is detected and the location of the closest manifold point yields the feature parameters. The concepts of parameter reduction by normalization, dimension reduction, pattern rejection, and heuristic search are all employed to achieve the required efficiency. By applying the algorithm to appropriate parametric feature models, detectors have been constructed for five features, namely, step edge, roof edge, line, corner, and circular disc. Detailed experiments are reported on the robustness of detection and the accuracy of parameter estimation.
Shree K. Nayar, Simon Baker, Hiroshi Murase
CVPR3
1996 Learning by a generation approach to appearance-based object recognition
abstract
We propose a methodology for the generation of learning samples in appearance-based object recognition. In many practical situations, it is not easy to obtain a large number of learning samples. The proposed method learns object models from a large number of generated samples derived from a small number of actually observed images. The learning algorithm has two steps: 1) generation of a large number of images by image interpolation, or image deformation, and 2) compression of the large sample sets using parametric eigenspace representation. We compare our method with the previous methods that interpolate sample points in eigenspace, and show the performance of our method to be superior. Experiments were conducted for 432 image samples for 4 objects to demonstrate the effectiveness of the method.
Hiroshi Murase, Shree K. Nayar
ICPR1
1996 Object location using complementary color features: histogram and DCT
abstract
Color constitutes an important cue for recognizing and locating objects in complex scenes. Most of the existing techniques using colors employ only color histograms for object recognition and/or location. Color histograms are stable but not accurate. In this paper we study the complementary nature of color histogram and DCT coefficients with respect to accuracy and stability and develop a combined method using both histograms and discrete cosine transform (DCT) coefficients. The methods are experimentally evaluated. The combined method has higher stability and accuracy than using either feature alone.
Vasudevan V. Vinod, Hiroshi Murase
ICPR2
1996 Dimensionality of illumination in appearance matching
abstract
Appearance matching was recently demonstrated as a robust and efficient approach to 3D object recognition and pose estimation. Each object is represented as a continuous appearance manifold in a low-dimensional subspace parametrized by object pose and illumination direction. Here, the structural properties of appearance manifolds are analyzed with the aim of making appearance representation efficient in off-line computation, storage requirements, and online recognition time. In particular, the effect of illumination on the structure of the appearance manifold is studied. It is shown that for an ideal diffused surface of arbitrary texture, the appearance manifold is linear and three dimensional. This enables the construction of the entire illumination manifold from just three images of the object taken using linearly independent light sources. This result is shown to hold even for illumination by multiple light sources and for concave surfaces that exhibit inter-reflections. Finally, a simple but efficient algorithm is presented that uses just three manifold points for recognizing images taken under novel illuminations.
Shree K. Nayar, Hiroshi Murase
ICRA2
1996 Real-time 100 object recognition system
abstract
A real-time vision system is described that can recognize 100 complex three-dimensional objects. In contrast to traditional strategies that rely on object geometry and local image features, the present system is founded on the concept of appearance matching. Appearance manifolds of the 100 objects were automatically learned using a computer-controlled turntable. The entire learning process was completed in 1 day. A recognition loop has been implemented that performs scene change detection, image segmentation, region normalizations, and appearance matching, in less than 1 second. The hardware used by the recognition system includes no more than a CCD color camera and a workstation. The real-time capability and interactive nature of the system have allowed numerous observers to test its performance. To quantify performance, we have conducted controlled experiments on recognition and pose estimation. The recognition rate was found to be 100% and object pose was estimated with a mean absolute error of 2.02 degrees and standard deviation of 1.67 degrees.
Shree K. Nayar, Sameer A. Nene, Hiroshi Murase
ICRA3
1996 Moving object recognition in eigenspace representation: gait analysis and lip reading
Hiroshi Murase, Rie Sakai
Pattern Recognit. Lett.1
1996 Subspace methods for robot vision
abstract
In contrast to the traditional approach, visual recognition is formulated as one of matching appearance rather than shape. For any given robot vision task, all possible appearance variations define its visual workspace. A set of images is obtained by coarsely sampling the workspace. The image set is compressed to obtain a low-dimensional subspace, called the eigenspace, in which the visual workspace is represented as a continuous appearance manifold. Given an unknown input image, the recognition system first projects the image to eigenspace. The parameters of the vision task are recognized based on the exact location of the projection on the appearance manifold. An efficient algorithm for finding the closest manifold point is described. The proposed appearance representation has several applications in robot vision. As examples, a precise visual positioning system, a real-time visual tracking system, and a real-time temporal inspection system are described.
Shree K. Nayar, Sameer A. Nene, Hiroshi Murase
IEEE Trans. Robotics Autom.3
1995 Visual learning and recognition of 3-d objects from appearance
Hiroshi Murase, Shree K. Nayar
Int. J. Comput. Vis.1
1995 Partial eigenvalue decomposition for large image sets using run-length encoding
James B. Roseborough, Hiroshi Murase
Pattern Recognit.2
1994 Illumination planning for object recognition in structured environments
abstract
This paper addresses the problem of illumination planning for robust object recognition in structured environments. Given a set of objects, the goal is to determine the illumination for which the objects are most distinguishable in appearance from each other. For each object, a large number of images is automatically obtained by varying pose and illumination. Images of all objects, together, constitute the planning image set. The planning set is compressed using the Karhunen-Loeve transform to obtain a low-dimensional subspace. For any given illumination, objects are represented as parametrized manifolds in the subspace. The minimum distance between the manifolds of too objects represents the similarity between the objects in the correlation sense. The optimal illumination is therefore one that maximizes the shortest distance between object manifolds. Results produced by the illumination planner heave been used to enhance the performance of an object recognition system.>
Hiroshi Murase, Shree K. Nayar
CVPR1
1994 Learning, Positioning, and Tracking Visual Appearance
abstract
The problem of vision-based robot positioning and tracking is addressed. A general learning algorithm is presented for determining the mapping between robot position and object appearance. The robot is first moved through several displacements with respect to its desired position, and a large set of object images is acquired. This image set is compressed using principal component analysis to obtain a four-dimensional subspace. Variations in object images due to robot displacements are represented as a compact parametrized manifold in the subspace. While positioning or tracking, errors in end-effector coordinates are efficiently computed from a single brightness image using the parametric manifold representation. The learning component enables accurate visual control without any prior hand-eye calibration. Several experiments have been conducted to demonstrate the practical feasibility of the proposed positioning/tracking approach and its relevance to industrial applications.>
Shree K. Nayar, Hiroshi Murase, Sameer A. Nene
ICRA2
1994 Illumination Planning for Object Recognition Using Parametric Eigenspaces
abstract
Presents a novel approach to the problem of illumination planning for robust object recognition in structured environments. Given a set of objects, the goal is to determine the illumination for which the objects are most distinguishable in appearance from each other. Correlation is used as a measure of similarity between objects. For each object, a large number of images is automatically obtained by varying the pose and the illumination direction. Images of all objects together constitute the planning image set. The planning set is compressed using the Karhunen-Loeve transform to obtain a low-dimensional subspace, called the eigenspace. For each illumination direction, objects are represented as parametrized manifolds in the eigenspace. The minimum distance between the manifolds of two objects represents the similarity between the objects in the correlation sense. The optimal source direction is therefore the one that maximizes the shortest distance between the object manifolds. Several experiments have been conducted using real objects. The results produced by the illumination planner have been used to enhance the performance of an object recognition system.>
Hiroshi Murase, Shree K. Nayar
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Learning Object Models from Appearance
Hiroshi Murase, Shree K. Nayar
AAAI1
1993 Silhouette-based object recognition through curvature scale space
abstract
A complete and practical isolated-object recognition system has been developed which is very robust with respect to scale, position and orientation changes of the objects as well as noise and local deformations of shape due to perspective projection, segmentation errors and non-rigid material used in some objects. The system has been tested on a wide variety of 3-D objects with different shapes and surface properties. A light-box setup is used to obtain silhouette images which are segmented to obtain the physical boundaries of the objects which are classified as either convex or concave. Convex curves are recognized using their four high-scale curvature extrema points. Curvature scale space (CSS) representations are computed for concave curves. The CSS representation is a multi-scale organization of the natural invariant features of a curve. A three-stage coarse-to-fine matching algorithm quickly detects the correct object in each case.>
Farzin Mokhtarian, Hiroshi Murase
ICCV2
1992 Surface Shape Reconstruction of a Nonrigid Transport Object Using Refraction and Motion
abstract
The appearance of a pattern behind a transparent, moving object is distorted by refraction at the moving object's surface. An algorithm for reconstructing the surface shape of a nonrigid transparent object, such as water, from the apparent motion of the observed pattern is described. This algorithm is based on the optical and statistical analysis of the distortions. It consists of four steps: extraction of optical flow, averaging of each point trajectory obtained from the optical flow sequence, calculation of the surface normal using optical characteristics, and reconstruction of the surface. The algorithm is applied to both synthetic and real images to demonstrate its performance.>
Hiroshi Murase
IEEE Trans. Pattern Anal. Mach. Intell.1
1992 On-line handwriting recognition
abstract
For large-alphabet languages, like Japanese, handwriting input using an online recognition technique is essential for input accuracy and speed. However, there are serious problems that prevent high recognition accuracy of unconstrained handwriting. First, the thousands of ideographic Japanese characters of Chinese origin (called Kanji) can be written with wide variations in the number and order of strokes and significant shape distortions. Also, writing box-free recognition of characters is required to create a better man-machine interface. Intense research performed over the past 15 years to answer the most pressing recognition problems is described. Prototype systems are also described. The man-machine interfaces made possible by online handwriting recognition and anticipated advances in both hardware and software are discussed.>
Toru Wakahara, Hiroshi Murase, Kazumi Odaka
Proc. IEEE2
1991 One-Line Recognition System for Free-Format Handwritten Japanese Characters
abstract
This paper describes an on-line recognition system for free-format handwritten Japanese character strings which may contain characters with separated constituents or overlapping characters. The recognition method for the system, called candidate lattice method, conducts segmentation and recognition of individual character candidates, and applies linguistic information to determine the most probable character string in order to achieve high recognition rates. Special hardware designed to realize a real-time recognition system is also introduced. The method used on the special hardware attained a segmentation rate of 98.8% and an overall recognition rate of 98.7% for 105 samples.
Hiroshi Murase
Int. J. Pattern Recognit. Artif. Intell.1
1991 A lie group theoretic approach to the invariance problem in feature extraction and object recognition
James B. Cole, Hiroshi Murase, Seiichiro Naito
Pattern Recognit. Lett.2
1990 Surface shape reconstruction of an undulating transparent object
abstract
An algorithm is described for reconstructing the surface shape of a nonrigid transparent object, such as water, from the apparent motion of the observed pattern. This algorithm is based on the optical and statistical analysis of the distortions. It consists of the following parts: extraction of optical flow, averaging of each point trajectory obtained from the optical flow sequence, calculation of the surface normal using optical characteristics, and reconstruction of the surface. The algorithm is applied to synthetic and real images to demonstrate its performance.>
Hiroshi Murase
ICCV1
1988 Online recognition of free-format Japanese handwritings
abstract
An online recognition method, called the candidate lattice method, is described for free-format written Japanese character strings, which may contain characters with separated constituents or overlapping characters. The method conducts segmentation and recognition of individual character-candidates, and applies linguistic information to determine the most probable character string to achieve high recognition rates. Special hardware designed to realize a real-time recognition system is also introduced. The method, used on special hardware, attained a segmentation rate of 98.8% and an overall recognition rate of 98.7% for 105 samples.>
Hiroshi Murase
ICPR1
1986 Online hand-sketched figure recognition
Hiroshi Murase, Toru Wakahara
Pattern Recognit.1