Kazuhiro Hotta

dblp:54/748 · DBLP profile ↗
← Back
54ranked-venue papers
17as first author
18since 2021 · last 2025
0000-0002-5675-8713ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 26 · 11 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2025 Domain Generalization Using Category Information Independent of Domain Differences
Reiji Saito, Kazuhiro Hotta
ICPRAM2
2025 Recognition of Pitching Habits Using Multimodal Data of RGB Video and Skeleton
abstract
This paper tackles a fine-grained action recognition task that aims to identify pitch-type-specific habits in baseball pitching motions. This task is highly challenging because it requires distinguishing subtle differences among pitch types within the common motion of “throwing.” Although it scarcely addressed in previous studies, it holds great potential for practical applications such as sports performance analysis. We propose a novel multimodal approach that integrates skeleton and RGB modalities. First, the Segment Anything Model (SAM) is employed to remove background information from RGB videos, thereby mitigating background bias. Next, Deep Canonical Correlation Analysis (DCCA) is applied to align the skeleton and RGB features. By further decomposing the RGB features into structural and textural components, the proposed method enhances discriminative performance. Experiments conducted on a custom pitching dataset collected from three pitchers demonstrate that our method outperforms conventional approaches in classification accuracy and effectively captures subtle pitching habits.
Satoki Hidaka, Kazuhiro Hotta
ISM2
2025 Video Classification of Marchantia Polymorpha Using a Video Vision Transformer with Emphasized Channel Information
abstract
In recent years, deep learning has been used to analyze plant and cell images. One protein, MpPICALM-K, is localized at the base of flagella in Marchantia polymorpha (M.polymorpha) spermatozoids. MpPICALM-K is considered to be involved in the motility of the flagella that enable spermatozoids to swim. For the analysis of MpPICALM-K, video classification and visualization are used, and improving their performance can provide deeper biological insights. In this study, we propose a module called Dimensional-Fortes to improve Video Vision Transformer (ViViT). Dimensional-Fortes improved 6.15% in classification accuracy compared with the previous method (Integration-Net). Incorporating temporal information caused variations in the heatmaps, confirming its integration into the spatial visualization. These advancements are expected to provide stronger evidence that MpPICALM-K is involved in spermatozoid motility.
Haruhiko Murata, Naoki Minamino, Takashi Ueda, Yohei Kondo, Kazuhiro Hotta
ISM5
2024 Tracking Correction Method for Rapid and Random Protein Molecules Movement
Satoshi Kamiya, Keisuke Toida, Taka-aki Tsunoyama, Kazuhiro Hotta
ACCV (2)4
2024 Boundary Contrastive Learning for Label-Efficient Medical Image Segmentation
Satoshi Kamiya, Kota Yamashita, Kazuhiro Hotta
BMVC3
2024 Vision Transformer Interpretability via Prediction of Image Reflected Relevance Among Tokens
Kento Sago, Kazuhiro Hotta
ICPRAM2
2024 Visualization of the Basis for Decisions by Selecting Layers Based on Model's Predictions Using the Difference Between Two Networks
Takahiro Sannomiya, Kazuhiro Hotta
ICPRAM2
2024 Improvement of TransUNet Using Word Patches Created from Different Dataset
Ayato Takama, Satoshi Kamiya, Kazuhiro Hotta
ICPRAM3
2023 Single-Particle Tracking by Graph Transformer
abstract
The immune system has been studied extensively by increasing the demand for object tracking of particles. However, because researches on single-particle tracking (SPT) by machine learning have not progressed yet, currently there is a reliance on software analysis despite low accuracy. There are three problems with SPT. First, there are no differences in the feature of each molecule, so tracking by feature differences is not possible. Second, it is difficult to predict the direction of molecular motion because it is random. Third, the high density of molecules causes frequent ID switches. Therefore, we propose Particle Tracking via Graph Transformer (PTGT), which takes into account the relationships among molecules, to solve these problems.
Satoshi Kamiya, Kazuhiro Hotta, Taka-aki Tsunoyama, Akihiro Kusumi
ICASSP2
2023 Lite-HRNet Plus: Fast and Accurate Facial Landmark Detection
abstract
Facial landmark detection is an essential technology for driver status tracking and has been in demand for real-time estimations. As a landmark coordinate prediction, heatmap-based methods are known to achieve a high accuracy, and Lite-HRNet can achieve a fast estimation. However, with Lite-HRNet, the problem of a heavy computational cost of the fusion block, which connects feature maps with different resolutions, has yet to be solved. In addition, the strong output module used in HRNetV2 is not applied to Lite-HRNet. Given these problems, we propose a novel architecture called Lite-HRNet Plus. Lite-HRNet Plus achieves two improvements: a novel fusion block based on a channel attention and a novel output module with less computational intensity using multi-resolution feature maps. Through experiments conducted on two facial landmark datasets, we confirmed that Lite-HRNet Plus further improved the accuracy in comparison with conventional methods, and achieved a state-of-the-art accuracy with a computational complexity with the range of 10M FLOPs.
Sota Kato, Kazuhiro Hotta, Yuhki Hatakeyama, Yoshinori Konishi
ICIP2
2023 DeformableFormer for Classifying Endoscopic Ultrasound-Guided Fine-Needle Biopsy in Pancreatic Diseases
abstract
The purpose of this paper is to classify from an unstained image whether it is available for examination or not, and to exceed the accuracy of visual classification by specialist physicians by machine learning. Currently, Vision Transformer and MetaFormer based PoolFormer have shown high accuracy for image classification. However, the pancreatic tissue fragment is a part of the image and has a complex shape, so the Vision Transformer, which processes the entire image, and the Pool-Former, which uses localized but fixed Convolution and Pooling, cannot classify it well. To address the problem, we require localized and image-specific feature extraction depending on the shape of a target. Therefore, we propose DeformableFormer, which enables local and dynamic feature extraction depending on the shape of the classification target in each image. To evaluate our method, we classify two categories of pancreatic tissue fragments; available and unavailable for examination. We demonstrated that our method outperformed the accuracy by specialist physicians, ViT and Poolformer.
Taiji Kurami, Takuya Ishikawa, Kazuhiro Hotta
ISM3
2023 Enlarged Large Margin Loss for Imbalanced Classification
abstract
We propose a novel loss function for imbalanced classification. LDAM loss, which minimizes a margin-based generalization bound, is widely utilized for class-imbalanced image classification. Although, by using LDAM loss, it is possible to obtain large margins for the minority classes and small margins for the majority classes, the relevance to a large margin, which is included in the original softmax cross entropy loss, is not be clarified yet. In this study, we reconvert the formula of LDAM loss using the concept of the large margin softmax cross entropy loss based on the softplus function and confirm that LDAM loss includes a wider large margin than softmax cross entropy loss. Furthermore, we propose a novel Enlarged Large Margin (ELM) loss, which can further widen the large margin of LDAM loss. ELM loss utilizes the large margin for the maximum logit of the incorrect class in addition to the basic margin used in LDAM loss. Through experiments conducted on imbalanced CIFAR datasets and large-scale datasets with long-tailed distribution, we confirmed that classification accuracy was much improved compared with LDAM loss and conventional losses for imbalanced classification.
Sota Kato, Kazuhiro Hotta
SMC2
2022 Local Embedding for Axial Attention
abstract
Recently, the researches on deep neural networks using self-attention have been actively conducted and shown to be effective. However, self-attention requires a large amount of computational cost. Axial Attention reduces the computational complexity by factorizing 2D self-attention into two 1D self-attentions, but the visualization of activation region was found to be inaccurate compared to CNN. The inaccurate of activation region means that unrelated regions to true class are referred for classification. It will decrease the accuracy of segmentation, which is a pixel-by-pixel classification, in identifying object boundaries. In this study, we attempted to reduce the problem by introducing a Local Embedding Unit to Axial Attention. In addition, by improving the structure of Axial Attention, we were able to improve the accuracy while reducing the computational cost and the number of parameters. Our model achieved useful results on the classification of ImageNet and segmentation of the CamVid dataset.
Ryouichi Furukawa, Kazuhiro Hotta
ICIP2
2022 Predicting Human Behavior Using 3D Loop ResNet
abstract
In this research, we would like to predict human behavior from video images to help humans in the future medical and nursing care fields. As a preliminary step, we will predict human actions. As a baseline, we used the 3D ResNet, which can handle both temporal and spatial features and can identify human actions with high accuracy. To improve the feature representation of conventional 3D ResNet for small action, we propose a method using loop feature extraction based on Convolutional LSTM in a residual block. We conducted experiments using the UCF101, Kinetics-400, and HMDB-51 datasets. As a result, we confirmed that our method with the introduction of the loop mechanism can obtain higher accuracy than the conventional 3D ResNet. Next, we created our own data set for predicting human behavior and used it for evaluation. We were able to achieve high accuracy, indicating that it is possible to predict behavior to some extent.
Yoshiki Kakamu, Kazuhiro Hotta
ICPR2
2022 Reconstructed Student-Teacher and Discriminative Networks for Anomaly Detection
abstract
Anomaly detection is an important problem in computer vision; however, the scarcity of anomalous samples makes this task difficult. Thus, recent anomaly detection methods have used only “normal images” with no abnormal areas for training. In this work, a powerful anomaly detection method is proposed based on student-teacher feature pyramid matching (STPM), which consists of a student and teacher network. Generative models are another approach to anomaly detection. They reconstruct normal images from an input and compute the difference between the predicted normal and the input. Unfortunately, STPM does not have the ability to generate normal images. To improve the accuracy of STPM, this work uses a student network, as in generative models, to reconstruct normal features. This improves the accuracy; however, the anomaly maps for normal images are not clean because STPM does not use anomaly images for training, which decreases the accuracy of the image-level anomaly detection. To further improve accuracy, a discriminative network trained with pseudo-anomalies from anomaly maps is used in our method, which consists of two pairs of student-teacher networks and a discriminative network. The method displayed high accuracy on the MVTec anomaly detection dataset.
Shinji Yamada, Satoshi Kamiya, Kazuhiro Hotta
IROS3
2022 Cell image segmentation by using feedback and convolutional LSTM
abstract
Abstract Human brain is known to have a layered structure and perform not only feedforward process from lower layer to upper layer, but also feedback process from upper layer to lower layer. Neural network is a mathematical model of the function of neurons, and several models are proposed until now. Although neural network imitates the human brain, everyone uses only feedforward process and direct feedback process from upper layer to lower layer is not used in prediction process. Therefore, in this paper, we propose Feedback U-Net using convolutional LSTM. Our model is a segmentation model using convolutional LSTM and feedback process. The output of U-Net at the first round is fed back to the input, and our method re-considers the segmentation result at the second round. By using convolutional LSTM, the features are extracted well based on the features extracted at the first round. On both of the Drosophila cell image and Mouse cell image datasets, our model outperformed conventional U-Net which uses only feedforward process.
Eisuke Shibuya, Kazuhiro Hotta
Vis. Comput.2
2021 Localized Feature Aggregation Module for Semantic Segmentation
abstract
We propose a new information aggregation method which called "Localized Feature Aggregation Module" based on the similarity between the feature maps of an encoder and a decoder. The proposed method recovers positional information by emphasizing the similarity between decoder’s feature maps with superior semantic information and encoder’s feature maps with superior positional information. The proposed method can learn positional information more efficiently than conventional con-catenation in the U-net and attention U-net. Additionally, the proposed method also uses localized attention range to reduce the computational cost. Two innovations contributed to improve the segmentation accuracy with lower computational cost. By experiments on the Drosophila cell image dataset and COVID-19 image dataset, we confirmed that our method outperformed conventional methods.
Ryouichi Furukawa, Kazuhiro Hotta
SMC2
2021 Automatic Preprocessing and Ensemble Learning for Cell Segmentation with Low Quality
abstract
We propose an automatic preprocessing and ensemble learning for segmentation of cell images with low quality. It is difficult to capture cells with strong light. Therefore, the microscopic images of cells tend to have low image quality but these images are not good for semantic segmentation. Here we propose a method to translate an input image to the images that are easy to recognize by deep learning. The proposed method consists of two deep neural networks. The first network is the usual training for semantic segmentation, and penultimate feature maps of the first network are used as filters to translate an input image to the images that emphasize each class. This is the automatic preprocessing and translated cell images are easily classified. The input cell image with low quality is translated by the feature maps in the first network, and the translated images are fed into the second network for semantic segmentation. Since the outputs of the second network are multiple segmentation results, we conduct the weighted ensemble of those segmentation images. Two networks are trained by end-to-end manner, and we do not need to prepare images with high quality for the translation. We confirmed that our proposed method can translate cell images with low quality to the images that are easy to segment, and segmentation accuracy has improved using the weighted ensemble learning.
Sota Kato, Kazuhiro Hotta
SMC2
2018 Road Detection from Satellite Images by Improving U-Net with Difference of Features
Ryosuke Kamiya, Kazuhiro Hotta, Kazuo Oda, Satomi Kakuta
ICPRAM2
2018 Semantic Segmentation in Red Relief Image Map by UX-Net
Tomoya Komiyama, Kazuhiro Hotta, Kazuo Oda, Satomi Kakuta, Mikako Sano
ICPRAM2
2018 Segmentation of Lidar Intensity using Weighted Fusion based on Appropriate Region Size
Masaki Umemura, Kazuhiro Hotta, Hideki Nonaka, Kazuo Oda
ICPRAM2
2018 Mixture of counting CNNs
Shohei Kumagai, Kazuhiro Hotta, Takio Kurita 0001
Mach. Vis. Appl.2
2016 PLSNet: A simple network using Partial Least Squares regression for image classification
abstract
PCANet is a simple network using Principal Component Analysis (PCA) for image classification and obtained high accuracies on a variety of datasets. PCA projects explanatory variables on a subspace that the first component has the largest variance. On the other hand, Partial Least Squares (PLS) regression projects explanatory variables on a subspace that the first component has the largest covariance between explanatory and objective variables, and the objective variables are predicted from the subspace. If class labels are used as objective variables for PLS, the subspace is suitable for classification. Stacked PLS is a simple network using PLS for image classification and obtained high accuracy on the MNIST database. However, the performance of Stacked PLS was inferior to PCANet on the others. One of differences between Stacked PLS and PCANet is network architecture. In this paper, we combine the network architecture of PCANet with PLS and propose a new image classification method called PLSNet. It obtained higher accuracies than PCANet on the MNIST and the CIFAR-10 datasets. Furthermore, we change how to make filters for extracting features at the second convolution layer, and we call it Improved PLSNet. It obtained higher accuracies than PLSNet. In addition, we give it deeper network architecture, and we call it Deep Improved PLSNet. It obtained higher accuracies than Improved PLSNet.
Ryoma Hasegawa, Kazuhiro Hotta
ICPR2
2016 Human tracking in crowded scenes using target information at previous frames
abstract
Human tracking in crowded scenes is a challenging problem because of frequent occlusion and presence of the tracking in similar regions. In this paper, we propose an online human tracking method which can handle occlusion and targets with similar regions. Our method compares the target region with a surrounding region and targets with similar regions at current frame. In addition, we also compare the target region at current and previous frames. We reduce the probabilities of uncommon colors at current and previous frames thereby improving the tracking accuracy. The effectiveness of the proposed method has been demonstrated via comparison with state-of-the-art trackers on the PETS2009 dataset.
Hiromasa Takada, Kazuhiro Hotta, Pranam Janney
ICPR2
2014 Robust Human Detection to Pose and Occlusion Using Bag-of-Words
abstract
To understand the human action in still images, it is effective to detect the human region. However, since appearance of human is much different due to pose and occlusion, the detection is quite difficult. Here we propose robust human detection method to pose and occlusion using Bag-of-Words (BoW). In general, the location information is helpful in classification. When the human has occlusion and pose changes, the location information makes the feature vector inconsistent. By using BoW which ignores the location information, we can obtain the consistent feature representation even if the local feature appears in different location by pose changes. Furthermore, when the part of human is occluded, BoW can also construct the feature representation from only visible part. Thus, BoW makes feature representation robust to partial occlusion and pose changes. By the comparison with deformable part model (DPM), the effectiveness of our method is demonstrated.
Yuta Tani, Kazuhiro Hotta
ICPR2
2013 Unsupervised Light Spot Detection using Background Subtraction
Takaya Niwa, Kazuhiro Hotta
ICPRAM2
2013 Image Labeling using Integration of Local and Global Features
Takuto Omiya, Kazuhiro Hotta
ICPRAM2
2013 Action Recognition Using Effective Mask Patterns Selected from a Classificational Viewpoint
abstract
This paper presents action recognition using effective mask patterns selected from an classificational viewpoint. Cubic higher-order local auto-correlation (CHLAC) feature is robust to position changes of human actions in a video, and its effectiveness for action recognition was already shown. However, the mask patterns for extracting cubic higher-order local auto-correlation (CHLAC) features are fixed. In other words, the mask patterns are independent of action classes, and the features extracted from those mask patterns are not specialized for each action. Thus, we propose automatic creation of specialized mask patterns for each action. Our approach consists of 2 steps. First, mask patterns are created by clustering of local spatio-temporal regions in each action. However, unnecessary mask patterns such as same patterns and mask patterns with all 0 or 1 are included. Then we select the effective mask patterns for classification by feature selection techniques. Through experiments using the KTH dataset, the effectiveness of our method is shown.
Takumi Hayashi, Kazuhiro Hotta
ISM2
2012 Melanosome Tracking by Bayes Theorem and Estimation of Movable Region
Toshiaki Okabe, Kazuhiro Hotta
ICPRAM (2)2
2012 Counting and radius estimation of lipid droplet in intracellular images
abstract
Light spot counting and radius estimation in intracellular images is important for investigating the cause of clinical condition. Since light spots are counted manually now, we propose automatic counting and radius estimation methods by computer. To count light spots by computer, we realize it by face detection manner with Support Vector Machine. To estimate radius of light spot, we use circle fitting by robust statistic. The proposed method gives higher accuracy than ImageJ which is the generated light spot detection software in cell biology. The effectiveness of the radius estimating is also shown by experiments.
Shohei Kumagai, Kazuhiro Hotta
SMC2
2012 Local co-occurrence features in subspace obtained by KPCA of local blob visual words for scene classification
Kazuhiro Hotta
Pattern Recognit.1
2011 Local autocorrelation of similarities with subspaces for shift invariant scene classification
Kazuhiro Hotta
Pattern Recognit.1
2010 Scene Classification Using Local Co-occurrence Feature in Subspace Obtained by KPCA of Local Blob Visual Words
abstract
In recent years, scene classification based on local correlation of binarized projection lengths in subspace obtained by Kernel Principal Component Analysis (KPCA) of visual words was proposed and its effectiveness was shown. However, local correlation of 2 binary features becomes 1 only when both features are 1. In other cases, local correlation becomes 0. This discarded information. In this paper, all kinds of co-occurrence of 2 binary features are used. This is the first device of our method. The second device is local Blob visual words. Conventional method made visual words from an orientation histogram on each grid. However, it is too local information. We use orientation histograms in a local Blob on grid as a basic feature and develop local Blob visual words. The third device is norm normalization of each orientation histogram in a local Blob. By normalizing local norm, the similarity between corresponding orientation histogram is reflected in subspace by KPCA. By these 3 devices, the accuracy is achieved more than 84% which is higher than conventional methods.
Kazuhiro Hotta
ICPR1
2010 Local normalized linear summation kernel for fast and robust recognition
Kazuhiro Hotta
Pattern Recognit.1
2009 Scene classification based on local autocorrelation of similarities with subspaces
abstract
This paper presents a scene classification method based on local autocorrelation of similarities with subspaces. Although conventional methods used bag-of-visual words for scene classification, superior accuracy of Kernel Principal Component Analysis (KPCA) of visual words to bag-of-visual words was reported. Here we also use KPCA of visual words to extract rich information for classification. In the original paper, all local parts mapped into subspace were integrated by summation to be robust to the order, the number, and shift of local parts. This approach discarded the effective properties for scene classification such as the relation with neighboring regions. To use them, we use Local AutoCorrelation (LAC) feature of the similarities with subspaces (outputs of KPCA of visual words). The feature has both the relation with neighboring regions and the robustness to shift of objects. The proposed method is compared with conventional scene classification methods using the same database and protocol. We demonstrate that the proposed method outperforms conventional methods.
Kazuhiro Hotta
ICIP1
2009 Pose independent object classification from small number of training samples based on kernel principal component analysis of local parts
Kazuhiro Hotta
Image Vis. Comput.1
2009 View independent face detection based on horizontal rectangular features and accuracy improvement using combination kernel of various sizes
Kazuhiro Hotta
Pattern Recognit.1
2009 Adaptive weighting of local classifiers by particle filters for robust tracking
Kazuhiro Hotta
Pattern Recognit.1
2008 Automatic Particle Detection and Counting by One-Class SVM from Microscope Image
Hinata Kuba, Kazuhiro Hotta, Haruhisa Takahashi
ICONIP (2)2
2008 Asbestos Detection from Microscope Images Using Support Vector Random Field of Local Color Features
Yoshitaka Moriguchi, Kazuhiro Hotta, Haruhisa Takahashi
ICONIP (2)2
2008 An Asbestos Counting Method from Microscope Images of Building Materials Using Summation Kernel of Color and Shape
Atsuo Nomoto, Kazuhiro Hotta, Haruhisa Takahashi
ICONIP (2)2
2008 Non-linear feature extraction by linear PCA using local kernel
abstract
This paper presents how to extract non-linear features by linear PCA. KPCA is effective but the computational cost is the drawback. To realize both non-linearity and low computational cost simultaneously, the idea of local kernel is used. The mapped features of the polynomial kernel can be described explicitly. When input features are divided into some local features and the polynomial kernel is applied to each local features independently, the dimension of mapped features does not become so high. In addition, the inner product with all local mapped features corresponds to the local summation kernel. Thus, KPCA with the local summation kernel can be solved by linear PCA. The proposed approach is evaluated in object categorization problem which requires high non-linearity and computational cost. The proposed method gives much higher accuracy than linear PCA. The computational cost is lower than KPCA though the accuracy is slightly worse than KPCA.
Kazuhiro Hotta
ICPR1
2008 Scene Classification Based on Multi-resolution Orientation Histogram of Gabor Features
Kazuhiro Hotta
ICVS1
2008 Object Categorization Based on Kernel Principal Component Analysis of Visual Words
abstract
Many researchers are studying object categorization problem. It is reported that bag of keypoints approach which is based on local features without topological information is effective for object categorization. Conventional bag of keypoints approach selects the visual words by clustering and uses the similarity with each visual word as the features for classification. In this paper, we model the ensemble of visual words, and the similarities with ensemble of visual words not each visual word are used for classification. Kernel principal component analysis (KPCA) is used to model them and extract the information specialized for each category. The projection length in subspace is used as features for support vector machine (SVM). There are two reasons why we use KPCA to model the ensemble of visual words. The first reason is to model the non-linear variations induced by various kinds of visual words. The second reason is that KPCA of local features is robust to pose variations. The proposed method is evaluated using Caltech 101 database. We confirm that the proposed method is comparable with the state of the art methods without absolute position information.
Kazuhiro Hotta
WACV1
2008 Robust face recognition under partial occlusion based on support vector machine with local Gaussian summation kernel
Kazuhiro Hotta
Image Vis. Comput.1
2006 Support Vector Machine with Weighted Summation Kernel Obtained by Adaboost
abstract
This paper presents Support Vector Machine (SVM) with weighted summation kernel obtained by Adaboost. In recent years, SVM with local summation kernel is proposed to use local features in SVM effectively. However, the computational cost of original method is high because all local kernels are used. In general, the effective position and size are different for recognition task. However, original method applied local kernel with same size to all scalar features and all local kernels are integrated with equal weight. To improve the performance and reduce the computational cost, local kernels with various size are prepared at all positions of a recognition target and effective local kernels are selected by Adaboost. Since 1-Nearest Neighbor (1-NN) of the output of local Gaussian kernel at certain size and position is used as weak learner, a strong classifier obtained by Adaboost becomes a new weighted summation kernel specialized for given recognition task. The proposed method is applied to face detection task. We confirmed that the proposed weighted summation kernel gives better performance than original local summation kernel though the proposed method uses the smaller number of local kernels than local summation kernel.
Kazuhiro Hotta
AVSS1
2004 A robust face detector under partial occlusion
abstract
This paper presents a robust face detector under partial occlusion. In recent years, the effectiveness of support vector machines (SVM) to object detection has been reported. However, conventional methods apply one kernel to global features. Therefore, those methods are not robust to occlusion because global features are influenced easily by noise or occlusion. To overcome this problem, SVM with local kernels is proposed. It is used to realize a robust face detector under partial occlusion. The robustness of the proposed method under partial occlusion is shown by using occluded face images. The proposed method can detect faces wearing sunglasses or a scarf. It is also confirmed that the proposed method is superior to the conventional SVM with global kernel.
Kazuhiro Hotta
ICIP1
2002 Object Detection Method Based on Local Kernels and Automatic Kernel Selection by Kullback-Leibler Divergence
abstract
This paper presents a object detection method based on local kernels. The local kernels are arranged to all positions on recognition target and are selected automatically by using Kullback-Leibler divergence according to the recognition target. The proposed method is applied to pedestrian detection problem. The performance of the proposed method is evaluated by the experiment using MIT CBCL pedestrian database. It is confirmed that generalization ability of the proposed method is improved by selecting the local kernels automatically.
Kazuhiro Hotta
WACV1
2000 Face Matching through Information Theoretical Attention Points and Its Applications to Face Detection and Classification
abstract
This paper presents a face matching method through information theoretical attention points. The attention points are selected as the points where the outputs of Gabor filters applied to the contrast-filtered image (Gabor features) have rich information. The information value of Gabor features of the certain point is used as the weight and the weighed sum of the correlations is used as the similarity measure for the matching. To cope with the scale changes of a face, several images with different scales are generated by interpolation from the input image and the best match is searched. By using the attention points given from the information theoretical point of view, the matching becomes robust under various environments. This matching method is applied to face detection of a known person and face classification. The effectiveness of the proposed method is confirmed by experiments using the face images captured over years under the different environments.
Kazuhiro Hotta, Taketoshi Mishima, Takio Kurita 0001, Shinji Umeyama
FG1
2000 Efficient Face Detection from News Images by Adaptive Estimation of Prior Probabilities and Ising Search
abstract
This paper presents an efficient method to detect faces from image sequences such as news or visual surveillance. To speed up the search, the prior probabilities of the face locations in the image are adaptively estimated and are used in the Ising search algorithm. The search points in the Ising search are selected depending on the estimated prior probabilities. The information obtained by the previous search points in the given image is effectively utilized through spin flip dynamics of the Ising search. If a face is found, the prior probabilities are updated with forgetting. This makes adaptation to the changes of the environment possible. The proposed search method was applied to images captured in the news. The proposed method is about 10 times faster than Ising search method without prior probabilities estimation.
Takio Kurita, Masaru Tanaka, Kazuhiro Hotta, Hiroyuki Shimai, Taketoshi Mishima
ICPR3
1999 Communicative functions to support human robot cooperation
abstract
We have been developing an autonomous robotic agent that helps people in a real world environment, such as in an office. When a robotic agent works by cooperating with a person in a real world environment, it must manage a lot of information and deal with the knowledge and languages that people usually use. Therefore it is important for the agent to recognize what people request as soon as possible. To realize common communication with people, the agent should provide robust communicative functions to obtain information from people. We discuss communicative functions of our robotic agent called Jijo-2. Especially we focus on the problem of detecting human faces, and discuss how a method of detecting a human face can be robustly archived.
Isao Hara, Alexander Zelinsky, Toshihiro Matsui, Hideki Asoh, Takio Kurita, Masaru Tanaka, Kazuhiro Hotta
IROS7
1998 Scale and Rotation Invariant Recognition Method Using Higher-Order Local Autocorrelation Features of Log-Polar Image
Takio Kurita 0001, Kazuhiro Hotta, Taketoshi Mishima
ACCV (2)2
1998 Scale Invariant Face Detection Method Using Higher-Order Local Autocorrelation Features Extracted from Log-Polar Image
Kazuhiro Hotta, Takio Kurita, Taketoshi Mishima
FG1
1998 Dynamic attention map by Ising model for human face detection
abstract
We present a method to narrow down the search space for scale-invariant human face detection, which uses a dynamic attention map implemented by Ising dynamics. Combining the proposed method and the scale-invariant face detection method which is based on both higher-order local autocorrelation (HLAC) features of a log-polar image and linear discriminant analysis for "face" and "not face" classification, it is shown that the "face" region in the image can be detected faster with some experiments.
Masaru Tanaka, Kazuhiro Hotta, Taketoshi Mishima, Takio Kurita
ICPR2