VLDB 2026 Research / reviewers in the wild / expert
Kazuhiro Fukui
dblp:01/1485
· DBLP profile ↗
66ranked-venue papers
6as first author
21since 2021 · last 2025
0000-0002-4201-1096ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Vision Language Model Interpretability with Concept Guided DecodingabstractVision Language Models (VLMs) are challenging for the field of Deep Learning Interpretability given their billions of parameters and recurrence-based reasoning processes. Based on the Frame Representation Hypothesis from language models, we introduce a novel approach for interpretability for the visual domain, enabling systematic extraction and analysis of learned concepts. Our framework reveals inherent biases and vulnerabilities in VLMs, caused by lack of proper safety alignment between the visual and textual components. By combining our interpretability techniques with the known FigStep jailbreak method, we demonstrate model vulnerabilities, achieving an average 97.7% attack success rate in the Qwen2-VL and Llama 3.2 Vision models. This combination of interpretability analysis and security testing provides crucial insights for developing safer, more transparent models, offering a path toward more trustworthy visual AI systems. Pedro H. V. Valois, Dipesh Satav, Rodrigo A. P. de Campos, Gulpi Qorik Oktagalu Pratamasunu, Kazuhiro Fukui |
ICIP | 5 |
| 2025 | Frame Representation Hypothesis: Multi-Token LLM Interpretability and Concept-Guided Text GenerationabstractAbstract Interpretability is a key challenge in fostering trust for Large Language Models (LLMs), which stems from the complexity of extracting reasoning from a model’s parameters. We present the Frame Representation Hypothesis, a theoretically robust framework grounded in the Linear Representation Hypothesis (LRH) to interpret and control LLMs by modeling multi-token words. Prior research explored LRH to connect LLM representations with linguistic concepts, but was limited to single token analysis. As most words are composed of several tokens, we extend LRH to multi-token words, thereby enabling usage on any textual data with thousands of concepts. To this end, we propose that words can be interpreted as frames, ordered sequences of vectors that better capture token-word relationships. Then, concepts can be represented as the average of word frames sharing a common concept. We showcase these tools through Top-k Concept-Guided Decoding, which can intuitively steer text generation using concepts of choice. We verify said ideas on Llama 3, Gemma 2, Phi 3, and Qwen-2-VL families, demonstrating gender and language biases, exposing harmful content, but also potential to remediate them, leading to safer and more transparent LLMs. Code is available at this https url. Pedro H. V. Valois, Lincon Sales de Souza, Erica K. Shimomoto, Kazuhiro Fukui |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Neural Network Innovations in Image-Based Malware Classification: A Comparative Study
Hamzah Al-Qadasi, Djafer Yahia Messaoud Benchadi, Salim Chehida, Kazuhiro Fukui, Saddek Bensalem |
AINA (4) | 4 |
| 2024 | Training-Free Zero-Shot Semantic Segmentation with LLM Refinement
Yuantian Huang, Satoshi Iizuka, Kazuhiro Fukui |
BMVC | 3 |
| 2024 | Sign Language Recognition Based on Subspace Representations in the Spatio-Temporal Frequency Domain
Ryota Sato, Suzana R. A. Beleza, Erica K. Shimomoto, Matheus Silva de Lima, Nobuko Kato, Kazuhiro Fukui |
ICPRAM | 6 |
| 2024 | Occlusion Sensitivity Analysis with Augmentation Subspace Perturbation in Deep Feature SpaceabstractDeep Learning of neural networks has gained prominence in multiple life-critical applications like medical diagnoses and autonomous vehicle accident investigations. However, concerns about model transparency and biases persist. Explainable methods are viewed as the solution to address these challenges. In this study, we introduce the Occlusion Sensitivity Analysis with Deep Feature Augmentation Subspace (OSA-DAS), a novel perturbation-based interpretability approach for computer vision. While traditional perturbation methods make only use of occlusions to explain the model predictions, OSA-DAS extends standard occlusion sensitivity analysis by enabling the integration with diverse image augmentations. Distinctly, our method utilizes the output vector of a DNN to build low-dimensional subspaces within the deep feature vector space, offering a more precise explanation of the model prediction. The structural similarity between these subspaces encompasses the influence of diverse augmentations and occlusions. We test extensively on the ImageNet-1k, and our class- and model-agnostic approach outperforms commonly used interpreters, setting it apart in the realm of explainable AI. Pedro H. V. Valois, Koichiro Niinuma, Kazuhiro Fukui |
WACV | 3 |
| 2024 | Controllable multi-domain semantic artwork synthesisabstractWe present a novel framework for the multi-domain synthesis of artworks from semantic layouts. One of the main limitations of this challenging task is the lack of publicly available segmentation datasets for art synthesis. To address this problem, we propose a dataset called ArtSem that contains 40,000 images of artwork from four different domains, with their corresponding semantic label maps. We first extracted semantic maps from landscape photography and used a conditional generative adversarial network (GAN)-based approach for generating high-quality artwork from semantic maps without requiring paired training data. Furthermore, we propose an artwork-synthesis model using domain-dependent variational encoders for high-quality multi-domain synthesis. Subsequently, the model was improved and complemented with a simple but effective normalization method based on jointly normalizing semantics and style, which we call spatially style-adaptive normalization (SSTAN). Compared to the previous methods, which only take semantic layout as the input, our model jointly learns style and semantic information representation, improving the generation quality of artistic images. These results indicate that our model learned to separate the domains in the latent space. Thus, we can perform fine-grained control of the synthesized artwork by identifying hyperplanes that separate the different domains. Moreover, by combining the proposed dataset and approach, we generated user-controllable artworks of higher quality than that of existing approaches, as corroborated by quantitative metrics and a user study. Yuantian Huang, Satoshi Iizuka, Edgar Simo-Serra, Kazuhiro Fukui |
Comput. Vis. Media | 4 |
| 2023 | Domain-Sum Feature Transformation For Multi-Target Domain Adaptation
Takumi Kobayashi 0001, Lincon Sales de Souza, Kazuhiro Fukui |
BMVC | 3 |
| 2023 | Diffusion-Based Semantic Image Synthesis from Sparse Layouts
Yuantian Huang, Satoshi Iizuka, Kazuhiro Fukui |
CGI | 3 |
| 2023 | Free-View Expressive Talking Head Video EditingabstractWe present a novel framework for talking head video editing, allowing users to freely edit head pose, emotion, and eye blink while maintaining audio-visual synchronization. Unlike previous approaches that mainly focus on generating a talking head video, our proposed model is able to edit the talking heads of an input video and restore it to full frames, which supports a broader range of applications. Our proposed framework consists of two parts: a) a reconstruction-based generator that can generate talking heads fitting to the original frame while corresponding to freely controllable attributes, including head pose, emotion, and eye blink. b) a multiple-attribute discriminator that enforces attribute-visual synchronization. We additionally introduce attention modules and perceptual loss to improve the overall generation quality. We compare existing approaches as corroborated by quantitative metrics and qualitative comparisons. Yuantian Huang, Satoshi Iizuka, Kazuhiro Fukui |
ICASSP | 3 |
| 2023 | Point Cloud Novelty Detection Based on Latent Representations of a General Feature Extractor
Shizuka Akahori, Satoshi Iizuka, Ken Mawatari, Kazuhiro Fukui |
PSIVT | 4 |
| 2023 | Diffusion-based Holistic Texture Rectification and SynthesisabstractWe present a novel framework for rectifying occlusions and distortions in degraded texture samples from natural images. Traditional texture synthesis approaches focus on generating textures from pristine samples, which necessitate meticulous preparation by humans and are often unattainable in most natural images. These challenges stem from the frequent occlusions and distortions of texture samples in natural images due to obstructions and variations in object surface geometry. To address these issues, we propose a framework that synthesizes holistic textures from degraded samples in natural images, extending the applicability of exemplar-based texture synthesis techniques. Our framework utilizes a conditional Latent Diffusion Model (LDM) with a novel occlusion-aware latent transformer. This latent transformer not only effectively encodes texture features from partially-observed samples necessary for the generation process of the LDM, but also explicitly captures long-range dependencies in samples with large occlusions. To train our model, we introduce a method for generating synthetic data by applying geometric transformations and free-form mask generation to clean textures. Experimental results demonstrate that our framework significantly outperforms existing methods both quantitatively and quantitatively. Furthermore, we conduct comprehensive ablation studies to validate the different components of our proposed framework. Results are corroborated by a perceptual user study which highlights the efficiency of our proposed approach. Guoqing Hao, Satoshi Iizuka, Kensho Hara, Edgar Simo-Serra, Hirokatsu Kataoka, Kazuhiro Fukui |
SIGGRAPH Asia | 6 |
| 2023 | Visually explaining 3D-CNN predictions for video classification with an adaptive occlusion sensitivity analysisabstractThis paper proposes a method for visually explaining the decision-making process of 3D convolutional neural networks (CNN) with a temporal extension of occlusion sensitivity analysis. The key idea here is to occlude a specific volume of data by a 3D mask in an input 3D temporalspatial data space and then measure the change degree in the output score. The occluded volume data that produces a larger change degree is regarded as a more critical element for classification. However, while the occlusion sensitivity analysis is commonly used to analyze single image classification, it is not so straightforward to apply this idea to video classification as a simple fixed cuboid cannot deal with the motions. To this end, we adapt the shape of a 3D occlusion mask to complicated motions of target objects. Our flexible mask adaptation is performed by considering the temporal continuity and spatial co-occurrence of the optical flows extracted from the input video data. We further propose to approximate our method by using the first-order partial derivative of the score with respect to an input image to reduce its computational cost. We demonstrate the effectiveness of our method through various and extensive comparisons with the conventional methods in terms of the deletion/insertion metric and the pointing metric on the UCF101. The code is available at: https://github.com/uchiyama33/AOSA. Tomoki Uchiyama, Naoya Sogi, Koichiro Niinuma, Kazuhiro Fukui |
WACV | 4 |
| 2023 | Grassmannian learning mutual subspace method for image set recognition
Lincon Sales de Souza, Naoya Sogi, Bernardo Bentes Gatto, Takumi Kobayashi 0001, Kazuhiro Fukui |
Neurocomputing | 5 |
| 2023 | Discriminant Feature Extraction by Generalized Difference SubspaceabstractIn this paper, we reveal the discriminant capacity of orthogonal data projection onto the generalized difference subspace (GDS), both theoretically and experimentally. In our previous work, we demonstrated that the GDS projection works as a quasi-orthogonalization of class subspaces, which is an effective feature extraction for subspace based classifiers. Here, we further show that GDS projection also works as a discriminant feature extraction through a similar mechanism to the Fisher discriminant analysis (FDA). A direct proof of the connection between GDS projection and FDA is difficult due to the significant difference in their formulations. To circumvent the complication, we first introduce geometrical Fisher discriminant analysis (gFDA) based on a simplified Fisher criterion. It is derived from a heuristic yet practically plausible assumption: the direction of the sample mean vector of a class is largely aligned to the first principal component vector of the class, given that the principal component analysis (PCA) is applied without data centering. gFDA works stably even under few samples, bypassing the small sample size (SSS) problem of FDA. We then prove that gFDA is equivalent to GDS projection with a small correction term. This equivalence ensures GDS projection to inherit the discriminant ability from FDA via gFDA. Furthermore, we discuss two useful extensions of these methods, 1) a nonlinear extension by kernel trick, 2) a combination with CNN features. The equivalence and the effectiveness of the extensions have been verified through extensive experiments on the extended Yale B+, CMU face database, ALOI, ETH80, MNIST, and CIFAR10, mainly focusing on image recognition under small samples. Kazuhiro Fukui, Naoya Sogi, Takumi Kobayashi 0001, Jing-Hao Xue, Atsuto Maki |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Image-set based Classification using Multiple Pseudo-whitened Mutual Subspace Method
Osamu Yamaguchi, Kazuhiro Fukui |
ICPRAM | 2 |
| 2022 | Constrained mutual convex cone method for image set based recognition
Naoya Sogi, Rui Zhu 0006, Jing-Hao Xue, Kazuhiro Fukui |
Pattern Recognit. | 4 |
| 2021 | Skeleton-aware Text Image Super-Resolution
Shimon Nakaune, Satoshi Iizuka, Kazuhiro Fukui |
BMVC | 3 |
| 2021 | Tensor analysis with n-mode generalized difference subspace
Bernardo Bentes Gatto, Eulanda M. dos Santos, Alessandro L. Koerich, Kazuhiro Fukui, Waldir S. S. Júnior |
Expert Syst. Appl. | 4 |
| 2021 | Correction to: Fukunaga-Koontz Convolutional Network with Applications on Character Classification
Bernardo Bentes Gatto, Eulanda M. dos Santos, Kazuhiro Fukui, Waldir S. S. Júnior, Kenny V. dos Santos |
Neural Process. Lett. | 3 |
| 2021 | Text classification based on the word subspace representation
Erica K. Shimomoto, François Portet, Kazuhiro Fukui |
Pattern Anal. Appl. | 3 |
| 2020 | Image Harmonization with Attention-based Deep Feature Modulation
Guoqing Hao, Satoshi Iizuka, Kazuhiro Fukui |
BMVC | 3 |
| 2020 | Resolving Marker Pose Ambiguity by Robust Rotation Averaging with Clique Constraints*abstractPlanar markers are useful in robotics and computer vision for mapping and localisation. Given a detected marker in an image, a frequent task is to estimate the 6DOF pose of the marker relative to the camera, which is an instance of planar pose estimation (PPE). Although there are mature techniques, PPE suffers from a fundamental ambiguity problem, in that there can be more than one plausible pose solutions for a PPE instance. Especially when localisation of the marker corners is noisy, it is often difficult to disambiguate the pose solutions based on reprojection error alone. Previous methods choose between the possible solutions using a heuristic criterion, or simply ignore ambiguous markers.We propose to resolve the ambiguities by examining the consistencies of a set of markers across multiple views. Our specific contributions include a novel rotation averaging formulation that incorporates long-range dependencies between possible marker orientation solutions that arise from PPE ambiguities. We analyse the combinatorial complexity of the problem, and develop a novel lifted algorithm to effectively resolve marker pose ambiguities, without discarding any marker observations. Results on real and synthetic data show that our method is able to handle highly ambiguous inputs, and provides more accurate and/or complete marker-based mapping and localisation. Shin-Fang Ch'ng, Naoya Sogi, Pulak Purkait, Tat-Jun Chin, Kazuhiro Fukui |
ICRA | 5 |
| 2020 | Discriminative Singular Spectrum Analysis for Bioacoustic Classification
Bernardo Bentes Gatto, Eulanda M. dos Santos, Juan Gabriel Colonna, Naoya Sogi, Lincon Sales de Souza, Kazuhiro Fukui |
INTERSPEECH | 6 |
| 2020 | Deep learning for image super-resolution
Wenming Yang, Fei Zhou 0001, Rui Zhu 0006, Kazuhiro Fukui, Guijin Wang, Jing-Hao Xue |
Neurocomputing | 4 |
| 2020 | Fukunaga-Koontz Convolutional Network with Applications on Character ClassificationabstractAbstract Several convolutional neural network architectures have been proposed for handwritten character recognition. However, most of the conventional architectures demand large scale training data and long training time to obtain satisfactory results. These requirements prevent the use of these methods in a broader range of applications. As an alternative to cope with these problems, we present a new convolutional network for handwritten character recognition based on the Fukunaga–Koontz transform (FKT). Our approach lies in the assumption that Fukunaga–Koontz convolutional kernels can be efficiently learned from subspaces and directly employed to produce high discriminant features in a shallow network architecture. When representing image classes by subspaces, the within-class separability is reduced, since the subspaces form clusters in a low-dimensional space. To increase the between-class separability, we compute a discriminative space from the training subspaces using FKT. By learning convolutional kernels from subspaces, it is possible to extract representative and discriminative features from an image with only a few parameters. Another contribution of the proposed network is the use of pooling layers, which further improves its performance. The proposed method, called Fukunaga–Koontz Network (FKNet), is suitable for solving practical problems, especially when training and processing times are constraints. Four publicly available handwritten character datasets are employed to evaluate the advantages of FKNet. In addition, we demonstrate the flexibility of the proposed method by experiments on LFW dataset. Bernardo Bentes Gatto, Eulanda M. dos Santos, Kazuhiro Fukui, Waldir S. S. Júnior, Kenny V. dos Santos |
Neural Process. Lett. | 3 |
| 2020 | Enhanced Grassmann discriminant analysis with randomized time warping for motion recognition
Lincon Sales de Souza, Bernardo Bentes Gatto, Jing-Hao Xue, Kazuhiro Fukui |
Pattern Recognit. | 4 |
| 2020 | A Novel Separating Hyperplane Classification Framework to Unify Nearest-Class-Model Methods for High-Dimensional DataabstractIn this article, we establish a novel separating hyperplane classification (SHC) framework to unify three nearest-class-model methods for high-dimensional data: the nearest subspace method (NSM), the nearest convex hull method (NCHM), and the nearest convex cone method (NCCM). Nearest-class-model methods are an important paradigm for the classification of high-dimensional data. We first introduce the three nearest-class-model methods and then conduct dual analysis for theoretically investigating them, to understand deeply their underlying classification mechanisms. A new theorem for the dual analysis of NCCM is proposed in this article by discovering the relationship between a convex cone and its polar cone. We then establish the new SHC framework to unify the nearest-class-model methods based on the theoretical results. One important application of this new SHC framework is to help explain empirical classification results: why one class model has a better performance than others on certain data sets. Finally, we propose a new nearest-class-model method, the soft NCCM, under the novel SHC framework to solve the overlapping class model problem. For illustrative purposes, we empirically demonstrate the significance of our SHC framework and the soft NCCM through two types of typical real-world high-dimensional data: the spectroscopic data and the face image data. Rui Zhu 0006, Ziyu Wang 0003, Naoya Sogi, Kazuhiro Fukui, Jing-Hao Xue |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Classification of Bioacoustic Signals with Tangent Singular Spectrum AnalysisabstractAutomatic classification of bioacoustic signals is an essential tool in biology for laborious tasks such as environmental monitoring in areas of difficult access. A working system applied in the field must be able to run on small scale machines and make reasonable predictions from a small sample of data. Recently, a method called Grassmann singular spectrum analysis (GSSA) was introduced as the latest development in a line of research where bioacoustic signals are represented by subspaces. While this paradigm is compact and introduces a straightforward discriminant analysis for classification, it is based on a Grassmann kernel, which approximates the Grassmann manifold by a reproducing Hilbert kernel space, thus depending on a choice of a dictionary and not being able to capture the signals complexity from a small class sample. In this paper, we propose a method named tangent singular spectrum analysis (TSSA), which continues to exploit the advantages of subspace representation but does not rely on approximating the Grassmann manifold by a low-dimensional kernel. We formulate a discriminant analysis on a tangent space to the data sample mean, using the extrinsic coordinates of the manifold. The validity of TSSA is demonstrated through experiments on the Amazon rainforest Anuran dataset. Lincon Sales de Souza, Bernardo Bentes Gatto, Kazuhiro Fukui |
ICASSP | 3 |
| 2019 | MLSNet: Resource-Efficient Adaptive Inference with Multi-Level Segmentation NetworksabstractIn this paper, we propose a multi-level convolutional network for semantic segmentation. The advantage of our network is that it allows us to adaptively control the balance between the classification accuracy and the inference speed, depending on the limited computational resource and the complexity of a given task. Our network realizes such adaptive mechanism by introducing a hierarchically-connected decoder with multilevel classifiers. The low-level classifier is used for inference when the application prioritizes low computational cost over accuracy, and the high-level classifier is utilized when more accurate prediction is required. This switching can be automatically performed by specifying only a threshold for classification performance. Besides, to boost the lower-level classifiers, we incorporate a knowledge distillation mechanism into our network. The entire network is trained in an end-to-end fashion. Experiments on semantic segmentation datasets demonstrate that our model overperforms the conventional approaches. Shuhei Yokoo, Satoshi Iizuka, Kazuhiro Fukui |
ICIP | 3 |
| 2018 | Grassmann Singular Spectrum Analysis for Bioacoustics ClassificationabstractBioacoustic signal classification is a powerful tool for biologists, assisting in tasks such as environmental monitoring of biomes in areas of difficult access, and providing clues about the evolution and categorization of animals from the perspective of similarity of their bioacoustic mechanisms. The recently proposed mutual singular spectrum analysis (MSSA) introduced a novel bioacoustic signal representation based on subspaces, which is compact and requires no cost intensive preprocessing techniques (e.g. segmentation, noise reduction or syllable extraction). However, MSSA has no discriminant mechanism to separate classes, and it assumes that a class is composed of linear combinations of the reference signals, which in practice is unlikely, and impairs study of the individuals' signals among the same species. In this paper we propose an extension named Grassmann singular spectrum analysis (GSSA), which preserves the advantages of MSSA in addition to the following contributions: we assume that a class may be composed of a set of subspaces and we simplify bioacoustic signal subspace representation by mapping the subspaces onto a Grassmann manifold; and we offer a discriminant mechanism to separate the species in a classification task. We demonstrate the validity of GSSA through a classification experiment on a publicly available bioacoustic signals dataset. Lincon Sales de Souza, Bernardo Bentes Gatto, Kazuhiro Fukui |
ICASSP | 3 |
| 2018 | Action Recognition Method Based on Sets of Time Warped ARMA ModelsabstractIn this paper, we propose a novel method for recognizing human actions from sequential body skeleton data. Our method is based on ARMA (Autoregressive Mean Average) model, which is constructed from the matrix of 3D joint positions time-series. The intrinsic structure of an action can be compactly summarized by the observability matrix of the ARMA model. Since the column vectors of an observability matrix span a subspace, given two ARMA models, we can measure the similarity between them by the canonical angles between the corresponding subspaces. This framework based on subspace representation is useful for action recognition. However, it does not work well when handling various actions with different action speeds, since optimal row size of each observability matrix depends on the action speed. To address this limitation, we perform a random sampling operation to the row elements in each observability matrix, while preserving the order of the elements. By repeating this operation, we generate a set of various time-warped ARMA models with various local motion speeds. The essence of this idea is that a whole set of such time-warped ARMA models is invariant to the changes in action speed. Furthermore, to construct an effective classification framework, we applied Grassmann discriminant analysis to the time-warped ARMA models. The effectiveness of the proposed method is demonstrated through comparison experiments with state-of-the-art methods on two public datasets: MSR 3D action dataset and UT-Kinect dataset. Naoya Sogi, Kazuhiro Fukui |
ICPR | 2 |
| 2018 | Text Classification Based On Word Subspace With Term-FrequencyabstractText classification has become indispensable due to the rapid increase of text in digital form. Over the past three decades, efforts have been made to approach this task using various learning algorithms and statistical models based on bag-of-words (BOW) features. Despite its simple implementation, BOW features lack of semantic meaning representation. To solve this problem, neural networks started to be employed to learn word vectors, such as the word2vec. Word2vec embeds word semantic structure into vectors, where the angle between vectors indicates the meaningful similarity between words. To measure the similarity between texts, we propose the novel concept of word subspace, which can represent the intrinsic variability of features in a set of word vectors. Through this concept, it is possible to model text from word vectors while holding semantic information. To incorporate the word frequency directly in the subspace model, we further extend the word subspace to the term-frequency (TF) weighted word subspace. Based on these new concepts, text classification can be performed under the mutual subspace method (MSM) framework. The validity of our modeling is shown through experiments on the Reuters text database, comparing the results to various state-of-art algorithms. Erica K. Shimomoto, Lincon Sales de Souza, Bernardo Bentes Gatto, Kazuhiro Fukui |
IJCNN | 4 |
| 2018 | A Method Based on Convex Cone Model for Image-Set Classification With CNN FeaturesabstractIn this paper, we propose a method for image-set classification based on convex cone models, focusing on the effectiveness of convolutional neural network (CNN) features as its input. CNN feature has non-negative values when using the rectified linear unit as an activation function. This naturally leads us to model a set of CNN features by a convex cone and measure the geometrical similarity of convex cones in classification. To achieve this framework, we define sequentially multiple angles between two convex cones by repeating the alternating least square method, and then define the geometrical similarity between the cones by using the obtained angles. Moreover, to enhance our method, we introduce a discriminant space, which maximizes the between-class variance (gaps) and minimizes the within-class variance of the projected convex cones onto the discriminant space, like Fisher discriminant analysis. Finally, the classification is conducted by measuring the similarity between projected convex cones. The effectiveness of the proposed method is demonstrated through evaluation experiments on a private database of a multi-view hand shape dataset, and two public databases. Naoya Sogi, Taku Nakayama, Kazuhiro Fukui |
IJCNN | 3 |
| 2018 | Cone-based joint sparse modelling for hyperspectral image classification
Ziyu Wang 0003, Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue |
Signal Process. | 3 |
| 2018 | Structural Class Classification of 3D Protein Structure Based on Multi-View 2D ImagesabstractComputing similarity or dissimilarity between protein structures is an important task in structural biology. A conventional method to compute protein structure dissimilarity requires structural alignment of the proteins. However, defining one best alignment is difficult, especially when the structures are very different. In this paper, we propose a new similarity measure for protein structure comparisons using a set of multi-view 2D images of 3D protein structures. In this approach, each protein structure is represented by a subspace from the image set. The similarity between two protein structures is then characterized by the canonical angles between the two subspaces. The primary advantage of our method is that precise alignment is not needed. We employed Grassmann Discriminant Analysis (GDA) as the subspace-based learning in the classification framework. We applied our method for the classification problem of seven SCOP structural classes of protein 3D structures. The proposed method outperformed the k-nearest neighbor method (k-NN) based on conventional alignment-based methods CE, FATCAT, and TM-align. Our method was also applied to the classification of SCOP folds of membrane proteins, where the proposed method could recognize the fold HEM-binding four-helical bundle (f.21) much better than TM-Align. Chendra Hadi Suryanto, Hiroto Saigo, Kazuhiro Fukui |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2017 | Subspace-Based Convolutional Network for Handwritten Character RecognitionabstractIn recent years, several convolutional neural networks-based architectures (CNN) have been proposed for handwritten character recognition. However, most of the conventional proposed models demand large scale training data and long training time to compute the parameters and achieve satisfactory results. These requirements prevent the use of these methods in a wider range of applications. To solve these problems, we present a novel convolutional network for handwritten character recognition based on subspace method. Our approach lies on the assumption that convolutional kernels can be efficiently learned from subspaces and directly employed to produce high discriminant features in a CNN architecture. When representing each image class by subspaces, we decrease the inter-class similarity, since the subspaces form clusters in a low dimensional space. In order to enlarge the intra-class similarity, we estimate a discriminative space from the training subspaces. By learning convolutional kernels from subspaces we can obtain representative and discriminative information produced by the network with few parameters, creating a light weight network. This use of a flexible architecture and its straightforward implementation make the proposed method quite attractive in practical terms. Our experimental evaluation shows that we achieve competitive results compared to the state-of-the-art methods. Bernardo Bentes Gatto, Eulanda M. dos Santos, Kazuhiro Fukui |
ICDAR | 3 |
| 2017 | Three-dimensional Object Recognition via Subspace Representation on a Grassmann Manifold
Ryoma Yataka, Kazuhiro Fukui |
ICPRAM | 2 |
| 2017 | Building a discriminatively ordered subspace on the generating matrix to classify high-dimensional spectral data
Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue |
Inf. Sci. | 2 |
| 2017 | Matched Shrunken Cone Detector (MSCD): Bayesian Derivations and Case Studies for Hyperspectral Target DetectionabstractHyperspectral images (HSIs) possess non-negative properties for both hyperspectral signatures and abundance coefficients, which can be naturally modeled using cone-based representation. However, in hyperspectral target detection, cone-based methods are barely studied. In this paper, we propose a new regularized cone-based representation approach to hyperspectral target detection, as well as its two working models by incorporating into the cone representation l2-norm and l1-norm regularizations, respectively. We call the new approach the matched shrunken cone detector (MSCD). Also important, we provide principled derivations of the proposed MSCD from the Bayesian perspective: we show that MSCD can be derived by assuming a multivariate half-Gaussian distribution or a multivariate half-Laplace distribution as the prior distribution of the coefficients of the models. In the experimental studies, we compare the proposed MSCD with the subspace methods and the sparse representation-based methods for HSI target detection. Two real hyperspectral data sets are used for evaluating the detection performances on sub-pixel targets and full-pixel targets, respectively. Results show that the proposed MSCD can outperform other methods in both cases, demonstrating the competitiveness of the regularized cone-based representation. Ziyu Wang 0003, Rui Zhu 0006, Kazuhiro Fukui, Jing-Hao Xue |
IEEE Trans. Image Process. | 3 |
| 2016 | Randomized time warping for motion recognition
Chendra Hadi Suryanto, Jing-Hao Xue, Kazuhiro Fukui |
Image Vis. Comput. | 3 |
| 2015 | Personal Authentication Based on 3D Configuration of Micro-feature Points on Facial Surface
Takao Yoshinuma, Hideitsu Hino, Kazuhiro Fukui |
PSIVT | 3 |
| 2015 | Difference Subspace and Its Generalization for Subspace-Based MethodsabstractSubspace-based methods are known to provide a practical solution for image set-based object recognition. Based on the insight that local shape differences between objects offer a sensitive cue for recognition, this paper addresses the problem of extracting a subspace representing the difference components between class subspaces generated from each set of object images independently of each other. We first introduce the difference subspace (DS), a novel geometric concept between two subspaces as an extension of a difference vector between two vectors, and describe its effectiveness in analyzing shape differences. We then generalize it to the generalized difference subspace (GDS) for multi-class subspaces, and show the benefit of applying this to subspace and mutual subspace methods, in terms of recognition capability. Furthermore, we extend these methods to kernel DS (KDS) and kernel GDS (KGDS) by a nonlinear kernel mapping to deal with cases involving larger changes in viewing direction. In summary, the contributions of this paper are as follows: 1) a DS/KDS between two class subspaces characterizes shape differences between the two respectively corresponding objects, 2) the projection of an input vector onto a DS/KDS realizes selective visualization of shape differences between objects, and 3) the projection of an input vector or subspace onto a GDS/KGDS is extremely effective at extracting differences between multiple subspaces, and therefore improves object recognition performance. We demonstrate validity through shape analysis on synthetic and real images of 3D objects as well as extensive comparison of performance on classification tests with several related methods; we study the performance in face image classification on the Yale face database B+ and the CMU Multi-PIE database, and hand shape classification of multi-view images. Kazuhiro Fukui, Atsuto Maki |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Sensing Visual Attention by Sequential PatternsabstractA method for sensing human visual attention is proposed. The method is based on the analysis of sequential image patterns of faces and irises observed at regular time intervals. The basic concept is to represent the set of image patterns produced by the action of gazing at a certain area as a nonlinear subspace in a high-dimensional pattern vector space. Such a space is called an attention subspace. In this framework, an input subspace from an unknown action is classified into an attention subspace of gazing at a certain area or into attention subspaces of gazing at other areas (named non-attention subspaces) by measuring the canonical angles between the input subspace and pre-computed dictionary subspaces. To maintain performance even in the presence of head movement, two mechanisms are introduced: 1) the kernel orthogonal mutual subspace method, which is suitable for classifying sets of multiple images, and 2) a kernel function for considering the head position in addition to a kernel function for pixel values. The stable performance of the proposed method including situations with head movements is demonstrated through experiments. Yasuyuki Yamazaki, Hideitsu Hino, Kazuhiro Fukui |
ICPR | 3 |
| 2014 | HEp-2 cell classification using rotation invariant co-occurrence among local binary patterns
Ryusuke Nosaka, Kazuhiro Fukui |
Pattern Recognit. | 2 |
| 2012 | A Scheme of Fragment-Based Faceted Image Search
Takahiro Komamizu, Mariko Kamie, Kazuhiro Fukui, Toshiyuki Amagasa, Hiroyuki Kitagawa |
DEXA (2) | 3 |
| 2012 | Towards a simulation driven stereo vision system
Martin Peris, Sara Martull, Atsuto Maki, Yasuhiro Ohkawa, Kazuhiro Fukui |
ICPR | 5 |
| 2012 | Protein structure similarity based on multi-view images generated from 3D molecular visualization
Chendra Hadi Suryanto, Shukun Jiang, Kazuhiro Fukui |
ICPR | 3 |
| 2011 | Feature Extraction Based on Co-occurrence of Adjacent Local Binary Patterns
Ryusuke Nosaka, Yasuhiro Ohkawa, Kazuhiro Fukui |
PSIVT (2) | 3 |
| 2010 | 3D Object Recognition Based on Canonical Angles between Shape Subspaces
Yosuke Igarashi, Kazuhiro Fukui |
ACCV (4) | 2 |
| 2009 | Image-Set Based Face Recognition Using Boosted Global and Local Principal Angles
Kazuhiro Fukui, Nanning Zheng 0001 |
ACCV (1) | 2 |
| 2009 | Boosting Constrained Mutual Subspace Method for Robust Image-Set Based Object Recognition
Kazuhiro Fukui, Nanning Zheng 0001 |
IJCAI | 2 |
| 2009 | Subspace based least squares support vector machines for pattern classificationabstractIn this paper, we discuss subspace based least squares support vector machines (SSLS-SVMs), in which an input vector is classified into the class with the maximum similarity. Namely, we define the similarity measure for each class by the weighted sum of vectors called dictionaries and optimize the weights so that the margin between classes is optimized. Because the similarity measure is defined for each class, the similarity measure associated with a data sample needs to be the largest among all the similarity measures. Introducing slack variables we define these constraints by equality constraints. Then the proposed SSLS-SVMs is similar to LS-SVMs by all-at-once formulation. Because all-at-once formulation is inefficient, we also propose SSLS-SVMs by one-against-all formulation. We demonstrate the effectiveness of the proposed methods with the conventional method for two-class problems. Takuya Kitamura, Shigeo Abe, Kazuhiro Fukui |
IJCNN | 3 |
| 2009 | Subspace based linear programming support vector machinesabstractIn subspace methods, the subspace associated with a class is represented by a small number of vectors called dictionaries and using the dictionaries the similarity measure is defined and an input is classified into the class with the highest similarity. Usually, each dictionary is given an equal weight. But if subspaces of different classes overlap, the similarity measures for the overlapping regions will not give useful information for classification. In this paper, we propose optimizing the weights for the dictionaries using the idea of support vector machines (SVMs). Namely, first we map the input space into the empirical feature space, perform kernel principal component analysis (KPCA) for each class, and define a similarity measure. Then considering that the similarity measure corresponds to the hyperplane, we formulate the optimization problem as maximizing the margin between the class associated with the dictionaries and the remaining classes. The optimization problem results in all-at-once formulation of linear SVMs. We demonstrate the effectiveness of the proposed method with that of the conventional methods for two-class problems. Syogo Takeuchi, Takuya Kitamura, Shigeo Abe, Kazuhiro Fukui |
IJCNN | 4 |
| 2009 | Classification of Similar 3D Objects with Different Types of Features from Multi-view Images
Hitoshi Niigaki, Kazuhiro Fukui |
PSIVT | 2 |
| 2009 | Subspace-based support vector machines for pattern classification
Takuya Kitamura, Syogo Takeuchi, Shigeo Abe, Kazuhiro Fukui |
Neural Networks | 4 |
| 2008 | Multiple view based 3D object classification using ensemble learning of local subspacesabstractMultiple observation improves the performance of 3D object classification. However, since the distribution of feature vectors obtained from multiple view points have strong nonlinear structure, the kernel-based methods are often introduced with nonlinear mapping. By mapping feature vectors to a higher dimensional space, kernel-based methods transform the distribution to weaken its nonlinearity. Although they have been succeeded in many applications, their computation cost is large. Therefore we aim to construct a comparable method with the kernel-based methods without using nonlinear mapping. Firstly we attempt to approximate a distribution of feature vectors with multiple local subspaces. Secondly we combine local subspace approximation with ensemble learning algorithm to form a new classifier. We will demonstrate that our method can achieve comparable performance with kernel methods through evaluation experiments using multiple view images of 3D objects from a public data set. Jianing Wu, Kazuhiro Fukui |
ICPR | 2 |
| 2007 | The Kernel Orthogonal Mutual Subspace Method and Its Application to 3D Object Recognition
Kazuhiro Fukui, Osamu Yamaguchi |
ACCV (2) | 1 |
| 2006 | A Framework for 3D Object Recognition Using the Kernel Constrained Mutual Subspace Method
Kazuhiro Fukui, Björn Stenger, Osamu Yamaguchi |
ACCV (2) | 1 |
| 2004 | Ship identification in sequential ISAR imagery
Atsuto Maki, Kazuhiro Fukui |
Mach. Vis. Appl. | 2 |
| 2003 | Face Recognition Using Multi-viewpoint Patterns for Robot Vision
Kazuhiro Fukui, Osamu Yamaguchi |
ISRR | 1 |
| 2002 | Constructing Illumination Image Basis from Object Motion
Akiko Nakashima, Atsuto Maki, Kazuhiro Fukui |
ECCV (3) | 3 |
| 2002 | Pattern hashing - object recognition based on a distributed local appearance modelabstractThis paper proposes "Pattern Hashing" as a new scheme for object recognition by effectively introducing an appearance-based approach into the framework of a geometric feature-based approach. We compose multiple bases using a combination of arbitrary three interest points in the model object, compute the geometric invariant for similarity transformation for each basis, and apply a hash function to it. Each image patch consists of pixels which are near the basis vector. We divide the model object image into multiple partial image patches, and create various appearances on the hash table as a distributed local appearance model. In the recognition stage, fast model selection is efficiently executed by the hashing technique, and then appearance pattern matching and voting procedure extract the target object in the input image. Through experiment with a face image database, we demonstrate that partly occluded object regions or multiple object positions can indeed be detected by the proposed algorithm. Osamu Yamaguchi, Kazuhiro Fukui |
ICIP (3) | 2 |
| 2000 | "GazeToTalk": a nonverbal interface with meta-communication facility (Poster Session)abstractWe propose a new human interface (HI) system named “GazeToTalk” that is implemented by vision based gaze detection, acoustic speech recognition (ASR), and animated human-like agent CG with facial expressions and gestures. The “GazeToTalk” system demonstrates that eye-tracking technologies can be utilized to improve HI effectively by working with other non-verbal messages such as facial expressions and gestures. Tetsuro Chino, Kazuhiro Fukui, Kaoru Suzuki |
ETRA | 2 |
| 1998 | Face Recognition Using Temporal Image Sequence
Osamu Yamaguchi, Kazuhiro Fukui, Ken-ichi Maeda |
FG | 2 |
| 1992 | Multiple object tracking system with three level continuous processesabstractReports a system for detecting human like moving objects in time-varying images. The authors show how it is possible to detect the image trajectories of people moving in ordinary indoor scenes. The system consists of three subprocesses: changing region detection, moving object tracking and movement interpretation. The processes are executed in parallel so hat each one can recover from the others' errors. This ensures the reliable detection of the trajectories in difficult cases such as movement across complicated backgrounds. The authors have built a trial detection system using a parallel image processing system. The details of the trial system and experimental results of walking person detection are described.> Kazuhiro Fukui, Hiroaki Nakai, Yoshinori Kuno |
WACV | 1 |