Keiichi Funaki

dblp:68/2772 · DBLP profile ↗
← Back
14ranked-venue papers
13as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 first-authorArtificial intelligence and machine learning · 8 · 7 first-authorSystems, architecture and hardware · 4 · 4 first-author · 3 since 2021
YearPublicationVenuePosition
2025 TLS-based TV-CAR speech analysis and evaluation of F0 estimation
abstract
This paper investigates Time-Varying Complex Auto Regressive (TV-CAR) speech analysis for analytic signals and its applications in speech processing. Our previous work has focused on developing an ℓ2-norm regularization scheme and a method based on Partial Least Squares (PLS). Additionally, we explore the Total Least Squares (TLS) method, which addresses linear equations with errors in both the observation matrix and the observation vector. Golub and Van Loan conducted extensive analysis of TLS from a numerical analysis perspective in the 1980s, leading to its widespread application in signal processing tasks, including spectral analysis, parameter estimation, adaptive filtering, and system identification. This paper proposes a TLS-based TV-CAR analysis method and evaluates its performance through fundamental frequency (F0) estimation of speech signals.
Keiichi Funaki
IECON1
2024 F0 estimation of speech using a complex adaptive pre-emphasis based on the ℓ2-norm regularized TV-CAR speech analysis
abstract
Linear Predictive (LP) analysis has been one of the most successful methods for speech analysis. It is widely used in applications such as smartphone speech coding and communication platforms like TEAMS and Zoom. However, LP analysis does have certain drawbacks, notably lower estimation accuracy due to its least-squares scheme. In response to these limitations, we have introduced an innovative approach called Time-Varying Complex Auto-Regressive (TV-CAR) speech analysis, which is built on Minimum Mean Square Error (MMSE), robustness, and LASSO analysis principles. We evaluated the performance of TV-CAR analysis regarding F0estimation in speech and robust speech recognition. Additionally, we proposed ℓ2-norm regularized TV-CAR analysis, Regularized LP (RLP)-based analysis, Time-RLP (TRLP)-based analysis, and hybrid methods that combine these approaches. We conducted evaluations focusing on F0estimation using the IRAPT method, leveraging complex residuals. Moreover, we previously introduced a Bone-Conducted (BC) prefilter to enhance overall performance. This paper presents an improved F0estimation technique that incorporates complex-valued Adaptive Pre-Emphasis (APE), showcasing its effectiveness. Furthermore, we delve into the frame-based computation of analytic signals.
Keiichi Funaki
IECON1
2021 On an improved F0 estimation based on ℓ2-norm regularized TV-CAR speech analysis using pre-filter
abstract
LP (Linear Prediction) analysis is the most commonly used and successful speech analysis implemented in smartphones. We have been proposing time-varying complex AR (TV-CAR) speech analysis for an analytic signal to solve the shortcomings of the LP. Recently, we proposed ℓ2-norm regularization TV-CAR analysis that penalizes rapid spectral changes in time-domain and frequency-domain, called hybrid RLP (Regularized LP) and TRLP (Time-RLP) method, and we have already evaluated performance using F0estimation for noise corrupted speech. The IRAPT (Instantaneous RAPT) for the estimated complex residual realizes the F0estimation. This paper proposes an improved F0estimation introducing a pre-filter as pre-processing and evaluates the performance.
Keiichi Funaki
IECON1
2012 Evaluation of F0 estimation using ZFR based on time-varying speech analysis
abstract
We have proposed F0estimation based on time-varying complex AR (TV-CAR) speech analysis in which F0is estimated using an weighted auto-correlation function for complex-valued residual signal calculated by the estimated time-varying complex-valued parameter for analytic signal. On the other hand, Zero Frequency Resonance (ZFR) has been proposed and it has been reported that the ZFR can estimate more accurate F0. The ZFR employs Zero Frequency Filtering (ZFF) for Hilbert envelope(HE) of LP residual to emphasize the resonance at zero frequency. In this paper, the ZFR based on TV-CAR speech analysis is proposed to estimate more accurate F0. In the proposed method, the HE is calculated with complex LP residual estimated by the complex parameters for analytic signal. The ZFR signal is calculated from the HE. The ZFR signal is used for the weighted auto-correlation to estimate F0. We have conducted the evaluation of F0estimation using Keele Pitch database. The experimental results show that LP residual-based ZFR method performs best.
Keiichi Funaki, Takehito Higa
ISCAS1
2010 On evaluation of the f0 estimation based on time-varying complex speech analysis
Keiichi Funaki
INTERSPEECH1
2002 Improvement of the ELS-based time-varying complex speech analysis
Keiichi Funaki
INTERSPEECH1
2001 A time-varying complex AR speech analysis based on GLS and ELS method
abstract
We have already developed three kinds of time-varying complex AR (TV-CAR) parameter estimation algorithms for analytic speech signal, which are based on minimizing mean square error (MMSE), Huber's robust Mestimation and Instrumental Variable (IV) method. This paper presents novel robust TV-CAR model parameter estimation algorithms on the basis of a Generalized Least Square (GLS) and Extended Least Square (ELS) method, in which the equation error is modeled by complex AR model with white Gaussian input to whiten the equation error. The experiments with natural speech corrupted by white Gaussian demonstrate that the proposed methods achieve robust spectral estimation against additive white Gaussian.
Keiichi Funaki
INTERSPEECH1
2000 A time-varying complex speech analysis based on IV method
Keiichi Funaki
INTERSPEECH1
2000 A study on the pitch pattern of a singing voice synthesis system based on the cepstral method
Tomio Takara, Kazuto Izumi, Keiichi Funaki
INTERSPEECH3
1999 Recursive ARMAX speech analysis based on a glottal source model with phase compensation
Keiichi Funaki, Yoshikazu Miyanaga, Koji Tochinai
Signal Process.1
1998 On robust speech analysis based on time-varying complex AR model
Keiichi Funaki, Yoshikazu Miyanaga, Koji Tochinai
ICSLP1
1997 A time varying ARMAX speech modeling with phase compensation using glottal source model
abstract
This paper presents new speech analysis method based on a glottal-ARMAX (autoregressive and moving average exogenous) model with phase compensation. A glottal-ARMAX model consists of two kinds of inputs: glottal source model excitation and a white Gaussian input, and a vocal tract ARMAX model. The proposed method can simultaneously estimate the glottal source model and vocal tract ARMAX model parameters pitch synchronously. In this method, ARMAX identification using a modified MIS (model identification system) method is adopted to estimate the ARMAX parameters, and the hybrid approach of the genetic algorithm (GA) and simulated annealing (SA) is employed to efficiently solve the non-linear simultaneous optimization of both parameters. Furthermore, phase compensation using an all-pass filter is introduced within a generation loop in the GA method in order to compensate phase distortion. Experiments using synthetic speech and natural speech demonstrate the efficacy of the proposed method.
Keiichi Funaki, Yoshikazu Miyanaga, Koji Tochinai
ICASSP1
1994 4kb/s speech coding with small computational amount and memory requirement: ULCELP
Keiichi Funaki, Kazunaga Yoshida, Kazunori Ozawa
ICSLP1
1990 A speech analysis method based on a glottal source model
Keiichi Funaki, Yukio Mitome
ICSLP1