EDBT 2026 Demo / reviewers in the wild / expert
Tomohiko Nakamura
dblp:04/5760
· DBLP profile ↗
25ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0003-4385-7170ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stride conversion algorithms for convolutional layers and its application to sampling-frequency-independent deep neural networksabstractWe propose interpolation-based algorithms that enable convolutional and transposed convolutional layers to operate with arbitrary (including non-integer) strides. A primary motivation for the proposed algorithms is to maintain a consistent temporal resolution when adapting deep neural networks (DNNs) to different sampling frequencies (SFs). To handle untrained SFs, we previously introduced SF-independent (SFI) convolutional layers, which adjust kernel weights in accordance with the target SF. However, achieving full consistency across SFs also requires the proportional adjustment of the stride, which results in non-integer values in many practical cases. Conventional algorithms for convolutional layers cannot handle such strides directly, and commonly used approaches (e.g., stride rounding or signal resampling) lead to performance degradation. To solve this problem, we propose a feature-domain interpolation framework that constructs continuous-time representations of intermediate features. This enables sampling at arbitrary stride intervals without modifying the network architecture. Through music source separation experiments, we show that the proposed algorithms maintain a strong performance across a range of SFs, including those where the stride becomes non-integer. Our analysis reveals that the proposed algorithms are robust to the choice of interpolation method and are especially effective for sources containing pitched sounds. Kanami Imamura, Tomohiko Nakamura, Norihiro Takamune, Kohei Yatabe, Hiroshi Saruwatari |
Signal Process. | 2 |
| 2025 | Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent LayerabstractWe introduce our submission to the AudioMOS Challenge (AMC) 2025 Track 3: mean opinion score (MOS) prediction for speech with multiple sampling frequencies (SFs). Our submitted model integrates an SF-independent (SFI) convolutional layer into a self-supervised learning (SSL) model to achieve SFI speech feature extraction for MOS prediction. We present some strategies to improve the MOS prediction performance of our model: distilling knowledge from a pretrained non-SFI-SSL model and pretraining with a large-scale MOS dataset. Our submission to the AMC 2025 Track 3 ranked the first in one evaluation metric and the fourth in the final ranking. We also report the results of our ablation study to investigate essential factors of our model. Go Nishikawa, Wataru Nakata, Yuki Saito 0001, Kanami Imamura, Hiroshi Saruwatari, Tomohiko Nakamura |
ASRU | 6 |
| 2024 | Neural Blind Source Separation and Diarization for Distant Speech Recognition
Yoshiaki Bando, Tomohiko Nakamura, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2024 | Self-Supervised Speech Representations are More Phonetic than Semantic
Kwanghee Choi, Ankita Pasad, Tomohiko Nakamura, Satoru Fukayama, Karen Livescu, Shinji Watanabe 0001 |
INTERSPEECH | 3 |
| 2024 | DNN-Based Ensemble Singing Voice Synthesis With Interactions Between SingersabstractWe propose a singing voice synthesis (SVS) method for a more unified ensemble singing voice by modeling interactions between singers. Most existing SVS methods aim to synthesize a solo voice, and do not consider interactions between singers, i.e., adjusting one’s own voice to the others’ voices. Since the production of ensemble voices from solo singing voices ignores the interactions, it can degrade the unity of the vocal ensemble. Therefore, we propose a SVS that reproduces the interactions. It is based on an architecture that uses musical scores of multiple voice parts, and loss functions that simulate the interactions’ effect to acoustic features. Experimental results show that our methods improve the unity of the vocal ensemble. Hiroaki Hyodo, Shinnosuke Takamichi, Tomohiko Nakamura, Junya Koguchi, Hiroshi Saruwatari |
SLT | 3 |
| 2023 | jaCappella Corpus: A Japanese a Cappella Vocal Ensemble CorpusabstractWe construct a corpus of Japanese a cappella vocal ensembles (ja-Cappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese children’s songs and have six voice parts (lead vocal, soprano, alto, tenor, bass, and vocal percussion). They are divided into seven subsets, each of which features typical characteristics of a music genre such as jazz and enka. The variety in genre and voice part match vocal ensembles recently widespread in social media services such as YouTube, although the main targets of conventional vocal ensemble datasets are choral singing made up of soprano, alto, tenor, and bass. Experimental evaluation demonstrates that our corpus is a challenging resource for vocal ensemble separation. Our corpus is available on our project page. Tomohiko Nakamura, Shinnosuke Takamichi, Naoko Tanji, Satoru Fukayama, Hiroshi Saruwatari |
ICASSP | 1 |
| 2023 | How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics
Joonyong Park, Shinnosuke Takamichi, Tomohiko Nakamura, Kentaro Seki, Detai Xin, Hiroshi Saruwatari |
INTERSPEECH | 3 |
| 2023 | TimToShape: Supporting Practice of Musical Instruments by Visualizing Timbre with 2D Shapes based on Crossmodal CorrespondencesabstractTimbre is high-dimensional and sensuous, making it difficult for musical-instrument learners to improve their timbre. Although some systems exist to improve timbre, they require expert labeling for timbre evaluation; however, solely visualizing the results of unsupervised learning lacks the intuitiveness of feedback because human perception is not considered. Therefore, we employ crossmodal correspondences for intuitive visualization of the timbre. We designed TimToShape, a system that visualizes timbre with 2D shapes based on the user’s input of timbre–shape correspondences. TimToShape generates a shape morphed by linear interpolation according to the timbre’s position in the latent space, which is obtained by unsupervised learning with a variational autoencoder (VAE). We confirmed that people perceived shapes generated by TimToShape to correspond more to timbre than randomly generated shapes. Furthermore, a user study of six violin players revealed that TimToShape was well-received in terms of visual clarity and interpretability. Kota Arai, Yutaro Hirao, Takuji Narumi, Tomohiko Nakamura, Shinnosuke Takamichi, Shigeo Yoshida |
IUI | 4 |
| 2023 | PoP-IDLMA: Product-of-Prior Independent Deeply Learned Matrix Analysis for Multichannel Music Source SeparationabstractIndependent deeply learned matrix analysis (IDLMA) is a state-of-the-art determined audio source separation method based on pretrained deep neural networks (DNNs). Owing to the excellent expression power of DNNs, IDLMA can handle a wider range of sources than conventional source models such as nonegative matrix factorization (NMF). However, owing to its supervised nature, the separation performance of IDLMA often degrades in the presence of timbral mismatches between the training data and the to-be-separated data. In this paper, we propose two source models that encompass the NMF- and DNN-based source models by constructing a prior distribution of the source power spectrogram (product of priors: PoP) on the basis of the product-of-expert concept. Since the NMF-based source model works well for a fully blind situation, the proposed models can handle the timbral mismatch without losing the expression power of DNNs. By introducing the PoP-based source models into IDLMA, we propose IDLMA extensions (PoP-IDLMAs) and derive their efficient parameter estimation algorithms on the basis of the majorization–minimization algorithm. Experimental results demonstrated the effectiveness of the proposed PoP-IDLMAs and that the proposed models greatly improve the source power estimation in frequency bands above 500 Hz. Takuya Hasumi, Tomohiko Nakamura, Norihiro Takamune, Hiroshi Saruwatari, Daichi Kitamura, Yu Takahashi, Kazunobu Kondo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Differentiable Digital Signal Processing Mixture Model for Synthesis Parameter Extraction from Mixture of Harmonic SoundsabstractA differentiable digital signal processing (DDSP) autoencoder is a musical sound synthesizer that combines a deep neural network (DNN) and spectral modeling synthesis. It allows us to flexibly edit sounds by changing the fundamental frequency, timbre feature, and loudness (synthesis parameters) extracted from an input sound. However, it is designed for a monophonic harmonic sound and cannot handle mixtures of harmonic sounds. In this paper, we propose a model (DDSP mixture model) that represents a mixture as the sum of the outputs of multiple pretrained DDSP autoencoders. By fitting the output of the proposed model to the observed mixture, we can directly estimate the synthesis parameters of each source. Through synthesis parameter extraction experiments, we show that the proposed method has high and stable performance compared with a straightforward method that applies the DDSP autoencoder to the signals separated by an audio source separation method. Masaya Kawamura, Tomohiko Nakamura, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo |
ICASSP | 2 |
| 2022 | SelfRemaster: Self-Supervised Speech Restoration with Analysis-by-Synthesis Approach Using Channel ModelingabstractWe present a self-supervised speech restoration method without paired speech corpora.Because the previous general speech restoration method uses artificial paired data created by applying various distortions to high-quality speech corpora, it cannot sufficiently represent acoustic distortions of real data, limiting the applicability.Our model consists of analysis, synthesis, and channel modules that simulate the recording process of degraded speech and is trained with real degraded speech data in a self-supervised manner.The analysis module extracts distortionless speech features and distortion features from degraded speech, while the synthesis module synthesizes the restored speech waveform, and the channel module adds distortions to the speech waveform.Our model also enables audio effect transfer, in which only acoustic distortions are extracted from degraded speech and added to arbitrary high-quality audio.Experimental evaluations with both simulated and real data show that our method achieves significantly higher-quality speech restoration than the previous supervised method, suggesting its applicability to real degraded speech materials. Takaaki Saeki, Shinnosuke Takamichi, Tomohiko Nakamura, Naoko Tanji, Hiroshi Saruwatari |
INTERSPEECH | 3 |
| 2022 | Sampling-Frequency-Independent Convolutional Layer and its Application to Audio Source SeparationabstractAudio source separation is often used for the preprocessing of various tasks, and one of its ultimate goals is to construct a single versatile preprocessor that can handle every variety of audio signal. One of the most important varieties of the discrete-time audio signal is sampling frequency. Since it is usually task-specific, the versatile preprocessor must handle all the sampling frequencies required by the possible downstream tasks. However, conventional models based on deep neural networks (DNNs) are not designed for handling a variety of sampling frequencies. Thus, for unseen sampling frequencies, they may not work appropriately. In this paper, we propose sampling-frequency-independent (SFI) convolutional layers capable of handling various sampling frequencies. The core idea of the proposed layers comes from our finding that a convolutional layer can be viewed as a collection of digital filters and inherently depends on sampling frequency. To overcome this dependency, we propose an SFI structure that features analog filters and generates weights of a convolutional layer from the analog filters. By utilizing time- and frequency-domain analog-to-digital filter conversion techniques, we can adapt the convolutional layer for various sampling frequencies. As an example application, we construct an SFI version of a conventional source separation network. Through music source separation experiments, we show that the proposed layers enable separation networks to consistently work well for unseen sampling frequencies in objective and perceptual separation qualities. We also demonstrate that the proposed method outperforms a conventional method based on signal resampling when the sampling frequencies of input signals are significantly lower than the trained sampling frequency. Koichi Saito, Tomohiko Nakamura, Kohei Yatabe, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Harmonic-Temporal Factor Decomposition for Unsupervised Monaural Separation of Harmonic SoundsabstractWe address the problem of separating a monaural mixture of harmonic sounds into the audio signals of individual semitones in an unsupervised manner. Unsupervised monaural audio source separation has thus far been mainly addressed by two approaches: one rooted in computational auditory scene analysis (CASA) and the other based on non-negative matrix factorization (NMF). These approaches focus on different clues for making source separation possible. A CASA-based method called harmonic-temporal clustering (HTC) focuses on a local time-frequency structure of individual sources, whereas NMF focuses on a global time-frequency structure of music spectrograms. These clues do not conflict with each other and can be used to achieve a more reliable audio source separation algorithm. Hence, we propose a monaural audio source separation framework, harmonic-temporal factor decomposition (HTFD), by developing a spectrogram model that encompasses the features of the models used in the NMF and HTC approaches. We further incorporate a source-filter model to build an extension of HTFD, source-filter HTFD (SF-HTFD). We derive efficient parameter estimation algorithms of HTFD and SF-HTFD based on the auxiliary function principle. We show, through music source separation experiments, the efficacy of HTFD and SF-HTFD compared with conventional methods. Furthermore, we demonstrate the effectiveness of HTFD and SF-HTFD for automatic musical key transposition. Tomohiko Nakamura, Hirokazu Kameoka |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Time-Domain Audio Source Separation With Neural Networks Based on Multiresolution AnalysisabstractWe propose a time-domain audio source separation method based on multiresolution analysis, which we call multiresolution deep layered analysis (MRDLA). The MRDLA model is based on one of the state-of-the-art time-domain deep neural networks (DNNs), Wave-U-Net, which successively down-samples features and up-samples them to have the original time resolution. From the signal processing viewpoint, we found that the down-sampling (DS) layers of Wave-U-Net cause aliasing and may discard information useful for source separation because they are implemented with decimation. These two problems are due to the decimation; thus, to achieve a more reliable source separation method, we should design DS layers capable of simultaneously overcoming these problems. With this motivation, focusing on the fact that the successive DS architecture of Wave-U-Net resembles that of multiresolution analysis, we develop DS layers based on discrete wavelet transforms (DWTs), which we call the DWT layers, because the DWTs have anti-aliasing filters and the perfect reconstruction property. We further extend the DWT layers such that their wavelet basis functions can be trained together with the other DNN components while maintaining the perfect reconstruction property. Since a straightforward trainable extension of the DWT layers does not guarantee the existence of anti-aliasing filters, we derive constraints for this guarantee in addition to the perfect reconstruction property. Through music source separation experiments including subjective evaluations, we show the efficacy of the proposed methods and the importance of simultaneously considering both the anti-aliasing filters and the perfect reconstruction property. Tomohiko Nakamura, Shihori Kozuka, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Time-Domain Audio Source Separation Based on Wave-U-Net Combined with Discrete Wavelet TransformabstractWe propose a time-domain audio source separation method using down-sampling (DS) and up-sampling (US) layers based on a discrete wavelet transform (DWT). The proposed method is based on one of the state-of-the-art deep neural networks, Wave-U-Net, which successively down-samples and up-samples feature maps. We find that this architecture resembles that of multiresolution analysis, and reveal that the DS layers of Wave-U-Net cause aliasing and may discard information useful for the separation. Although the effects of these problems may be reduced by training, to achieve a more reliable source separation method, we should design DS layers capable of overcoming the problems. With this belief, focusing on the fact that the DWT has an anti-aliasing filter and the perfect reconstruction property, we design the proposed layers. Experiments on music source separation show the efficacy of the proposed method and the importance of simultaneously considering the anti-aliasing filters and the perfect reconstruction property. Tomohiko Nakamura, Hiroshi Saruwatari |
ICASSP | 1 |
| 2016 | Shifted and convolutive source-filter non-negative matrix factorization for monaural audio source separationabstractThis paper proposes an extension of non-negative matrix factorization (NMF), which combines the shifted NMF model with the source-filter model. Shifted NMF was proposed as a powerful approach for monaural source separation and multiple fundamental frequency (F0) estimation, which is particularly unique in that it takes account of the constant inter-harmonic spacings of a harmonic structure in log-frequency representations and uses a shifted copy of a spectrum template to represent the spectra of different F0s. However, for those sounds that follow the source-filter model, this assumption does not hold in reality, since the filter spectra are usually invariant under F0 changes. A more reasonable way to represent the spectrum of a different F0 is to use a shifted copy of a harmonic structure template as the excitation spectrum and keep the filter spectrum fixed. Thus, we can describe the spectrogram of a mixture signal as the sum of the products between the shifted copies of excitation spectrum templates and filter spectrum templates. Furthermore, the time course of filter spectra represents the dynamics of the timbre, which is important for characterizing the feature of an instrument sound. Thus, we further incorporate the non-negative matrix factor deconvolution (NMFD) model into the above model to describe the filter spectrogram. We derive a computationally efficient and convergence-guaranteed algorithm for estimating the unknown parameters of the constructed model based on the auxiliary function approach. Experimental results revealed that the proposed method outperformed shifted NMF in terms of the source separation accuracy. Tomohiko Nakamura, Hirokazu Kameoka |
ICASSP | 1 |
| 2016 | Real-Time Audio-to-Score Alignment of Music Performances Containing Errors and Arbitrary Repeats and SkipsabstractThis paper discusses real-time alignment of audio signals of music performance to the corresponding score (a.k.a. score following) which can handle tempo changes, errors and arbitrary repeats and/or skips (repeats/skips) in performances. This type of score following is particularly useful in automatic accompaniment for practices and rehearsals, where errors and repeats/skips are often made. Simple extensions of the algorithms previously proposed in the literature are not applicable in these situations for scores of practical length due to the problem of large computational complexity. To cope with this problem, we present two hidden Markov models of monophonic performance with errors and arbitrary repeats/skips, and derive efficient score-following algorithms with an assumption that the prior probability distributions of score positions before and after repeats/skips are independent from each other. We confirmed real-time operation of the algorithms with music scores of practical length (around 10000 notes) on a modern laptop and their tracking ability to the input performance within 0.7 s on average after repeats/skips in clarinet performance data. Further improvements and extension for polyphonic signals are also discussed. Tomohiko Nakamura, Eita Nakamura, Shigeki Sagayama |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Lp-norm non-negative matrix factorization and its application to singing voice enhancementabstractMeasures of sparsity are useful in many aspects of audio signal processing including speech enhancement, audio coding and singing voice enhancement, and the well-known method for these applications is non-negative matrix factorization (NMF), which decomposes a non-negative data matrix into two non-negative matrices. Although previous studies on NMF have focused on the sparsity of the two matrices, the sparsity of reconstruction errors between a data matrix and the two matrices is also important, since designing the sparsity is equivalent to assuming the nature of the errors. We propose a new NMF technique, which we called Lp-norm NMF, that minimizes the Lpnorm of the reconstruction errors, and derive a computationally efficient algorithm for Lp-norm NMF according to an auxiliary function principle. This algorithm can be generalized for the factorization of a real-valued matrix into the product of two real-valued matrices. We apply the algorithm to singing voice enhancement and show that adequately selecting p improves the enhancement. Tomohiko Nakamura, Hirokazu Kameoka |
ICASSP | 1 |
| 2014 | Fast Signal Reconstruction from Magnitude Spectrogram of Continuous Wavelet Transform Based on Spectrogram Consistency
Tomohiko Nakamura, Hirokazu Kameoka |
DAFx | 1 |
| 2014 | Underdetermined blind separation and tracking of moving sources based ONDOA-HMMabstractThis paper deals with the problem of the underdetermined blind separation and tracking of moving sources. In practical situations, sound sources such as human speakers can move freely and so blind separation algorithms must be designed to track the temporal changes of the impulse responses. We propose solving this problem through the posterior inference of the parameters in a generative model of an observed multichannel signal, formulated under the assumption of the sparsity of time-frequency components of speech and the continuity of speakers' movements. Specifically, we describe a generative model of mixture signals by incorporating a generative model of a time-varying frequency array response for each source, described using a path-restricted hidden Markov model (HMM). Each hidden state of the present HMM represents the direction of arrival (DOA) of each source, and so we call it a “DOA-HMM.” Through the posterior inference of the overall generative model, we can simultaneously track the DOAs of sources, separate source signals and perform permutation alignment. The experiment showed that the proposed algorithm provided a 6.20 dB improvement compared with the conventional method in terms of the signal-to-interference ratio. Takuya Higuchi, Norihiro Takamune, Tomohiko Nakamura, Hirokazu Kameoka |
ICASSP | 3 |
| 2014 | Timbre replacement of harmonic and drum components for music audio signalsabstractThis paper presents a system that allows users to customize an audio signal of polyphonic music (input), without using musical scores, by replacing the frequency characteristics of harmonic sounds and the timbres of drum sounds with those of another audio signal of polyphonic music (reference). To develop the system, we first use a method that can separate the amplitude spectra of the input and reference signals into harmonic and percussive spectra. We characterize frequency characteristics of the harmonic spectra by two envelopes tracing spectral dips and peaks roughly, and the input harmonic spectra are modified such that their envelopes become similar to those of the reference harmonic spectra. The input and reference percussive spectrograms are further decomposed into those of individual drum instruments, and we replace the timbres of those drum instruments in the input piece with those in the reference piece. Through the subjective experiment, we show that our system can replace drum timbres and frequency characteristics adequately. Tomohiko Nakamura, Hirokazu Kameoka, Kazuyoshi Yoshii, Masataka Goto |
ICASSP | 1 |
| 2014 | A unified approach for underdetermined blind signal separation and source activity detection by multichannel factorial hidden Markov modelsabstractThis paper proposes to introduce a new model called “the multichannel factorial hidden Markov Model (MFHMM)” for underdetermined blind signal separation (BSS). For monaural source separation, one successful approach involves applying nonnegative matrix factorization (NMF) to the magnitude spectrogram of a mixture signal, interpreted as a non-negative matrix. Up to now, multichannel extensions of NMF, which allow for the use of spatial information as an additional clue for source separation, have been proposed by several authors and proven to be an effective approach for underdetermined BSS. This approach is based on the assumption that an observed signal is a mixture of a limited number of source signals each of which has a static power spectral density scaled by a time-varying amplitude. However, many source signals in real world are nonstationary in nature and the variations of the spectral densities are much richer in time. Moreover, many sources including speech tend to stay inactive for some while until they switch to an active mode, implying that the total power of a source may depend on its underlying state. To reasonably characterize such a non-stationary nature of source signals, this paper proposes to extend the multichannel NMF model by modeling the transition of the set consisting of the spectral densities and the total power of each source using a hidden Markov model (HMM). By letting each HMM contain states corresponding to active and inactive modes, we will show that voice activity detection and source separation can be solved simultaneously through parameter inference of the present model. The experiment showed that the proposed algorithm provided a 7.65 dB improvement compared with the conventional multichannel NMF in terms of the signal-to-distortion ratio. Index Terms: blind signal separation, source activity detection, a hidden Markov model, non-negative matrix factorization Takuya Higuchi, Hirofumi Takeda, Tomohiko Nakamura, Hirokazu Kameoka |
INTERSPEECH | 3 |
| 2004 | A design of category classification system for high resolution satelliteabstractFor the high resolution satellite image database obtained from IKONOS and Quickbird, we proposed a design including category classification system. These images are essentially different from the image NOAA AVHRR, Landsat TM, Spot, etc., because obtained image is understood like an aeronautical photograph. This system has a knowledge collection and high quality training function to classify the category. Masanori Nakano, Kazi A. Kalpoma, Tomohiko Nakamura, Jun-ichi Kudoh |
IGARSS | 3 |
| 2003 | Three dimensional histogram technique for IKONOS imagesabstractThe three dimensional histogram technique has been developed as a method for multi spectral image analysis such as NOAA AVHRR images. This report presents an approach of this method to classification for IKONOS satellite images. It is shown that four categories cropped from the image are separated in the three dimensional histogram. Jun-ichi Kudoh, Tomohiko Nakamura, S. Shikano, E. Kikuchi |
IGARSS | 2 |
| 1999 | Hierarchical-Clustering of Parametric Data with Application to the Parametric Eigenspace MethodabstractA novel method for hierarchical-clustering of parametric data is proposed (in this article, we assumed that these data were parameterized by a single parameter in multidimensional spaces). In the proposed clustering method, the continuity of a parameter is preserved in each class, and furthermore, the optimal loci of class boundaries and the appropriate number of classes are determined. To improve the recognition performance of the parametric eigenspace method, which is an object recognition method based on the visual learning approach, the proposed clustering method is applied to the construction of tree-structured dictionaries for this recognition method. Experimental result shows that these tree-structured dictionaries improve the recognition performance of the parametric eigenspace method without a decrease in recognition accuracy. Toru Abe, Tomohiko Nakamura |
ICIP (4) | 2 |