Yutaka Kamamoto

dblp:94/1145 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
3since 2021 · last 2025
0000-0001-6578-9178ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 4Databases, data management, data science and information retrieval · 4 · 1 first-authorTheory of computation · 3 · 2 since 2021
YearPublicationVenuePosition
2025 Stereo Downmix in 3GPP IVAS for EVS Compatibility
abstract
The 3GPP IVAS codec specifies an EVS-compatible stereo downmix as one of the key functionalities. This paper describes how this novel active downmix scheme has been devised to achieve high and stable quality from stereo input to EVS encoder/decoder with no additional algorithmic delay. An example of a network configuration for a multi-party conference using the proposed EVS-compatible downmix is also provided. This paper shows several fundamental schemes, of active processes, including adaptive weight and phase-compensated weight between two channels. Subjective listening test results show that the devised scheme’s quality is better than that of a passive downmix.
Takehiro Moriya, Stéphane Ragot, Arnaud Lefort, Alexandre Guérin, Noboru Harada, Ryosuke Sugiura, Yutaka Kamamoto
ICASSP7
2025 Optimal Construction of N-Bit-Delay Almost Instantaneous Fixed-to-Variable-Length Codes
abstract
This paper presents an optimal construction ofN-bit-delay almost instantaneous fixed-to-variable-length (AIFV) codes, the general form of binary codes we can make when finite bits of decoding delay are allowed. The presented method enables us to optimize lossless codes among a broader class of codes compared to the conventional FV and AIFV codes. The paper first discusses the problem of code construction, which contains some essential partial problems, and defines three classes of optimality to clarify how far we can solve the problems. The properties of the optimal codes are analyzed theoretically, showing the sufficient conditions for achieving the optimum. Then, we propose an algorithm for constructingN-bit-delay AIFV codes for given stationary memory-less sources. The optimality of the constructed codes is discussed both theoretically and empirically. They showed shorter expected code lengths whenN≥ 3 than the conventional AIFV-mand extended Huffman codes. Moreover, in the random numbers simulation, they performed higher compression efficiency than the 32-bit-precision range codes under reasonable conditions.
Ryosuke Sugiura, Masaaki Nishino, Norihito Yasuda, Yutaka Kamamoto, Takehiro Moriya
IEEE Trans. Inf. Theory4
2023 General Form of Almost Instantaneous Fixed-to-Variable-Length Codes
abstract
A general class of the almost instantaneous fixed-to-variable-length (AIFV) codes is proposed, which contains every possible binary code we can make when allowing finite bits of decoding delay. The contribution of the paper lies in the following. (i) Introducing$N$-bit-delay AIFV codes, constructed by multiple code trees with higher flexibility than the conventional AIFV codes. (ii) Proving that the proposed codes can represent any uniquely-encodable and uniquely-decodable variable-to-variable length codes. (iii) Showing how to express codes as multiple code trees with minimum decoding delay. (iv) Formulating the constraints of decodability as the comparison of intervals in the real number line. The theoretical results in this paper are expected to be useful for further study on AIFV codes.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
IEEE Trans. Inf. Theory2
2019 Shape Control of Discrete Generalized Gaussian Distributions for Frequency-Domain Audio Coding
abstract
Entropy coding, which is an essential part of audio compression, is always required to manage the tradeoffs between compression efficiency and computational complexity, and the strategy to achieve them highly depends on the distributions of inputs. In this paper, we present a method of controlling them for enhancing the compression efficiency of Golomb-Rice (GR) encoding, one of the simplest entropy coding methods optimal for Laplacian distributions. We will show that the proposed invertible and low-complexity mapping of integers enables the GR encoding to assign nearly the optimal code length for a wider range of distributions, generalized Gaussian distributions, maintaining low computational cost. A simulation by random numbers reveals that the proposed coder based on this scheme works about 6 times faster than the state-of-the-art arithmetic coder for Gaussian-distributed integers maintaining the increase in relative redundancy around 2.6%, which is much lower than that of a conventional GR coder. Additionally, an application to a practical speech and audio coding scheme is presented, and an objective evaluation for real speech and audio signals confirms the advantages of the proposed method in compression. The method is expected to widen the capability of low-complexity entropy coding, providing us with more flexible codec designs.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Integer Nesting/Splitting for Golomb-Rice Coding of Generalized Gaussian Sources
abstract
This paper presents a qualitative approach of combining Golomb-Rice (GR) code with algebraic bijective mappings which losslessly convert between arbitrary positive integers of different dimension and shape the distribution of generalized Gaussian sources. The mappings, integer nesting and splitting, enables GR encoding, with a little additional computation, to compress more efficiently sources based on wider classes of distributions than Laplacian. Simulations showed, especially for some Gaussian sources, almost optimal average code length can be achievable by performing integer nesting before GR encoding the integers. This scheme will be useful for applications dealing with various types of sources and requiring low computational costs.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
DCC2
2018 Spectral-Envelope-Based Least Significant Bit Management for Low-Delay Bit-Error-Robust Speech Coding
abstract
We have devised a method for bit assignment of quantized frequency spectra aiming at its use in low-delay bit-error-robust speech compression. The proposed method, least significant bit management (LSBM), controls the least significant bits of the spectra based on their envelope to make them represented by fixed bit rates, which guarantees the range of the damage caused by the bit error which makes the mismatch of the spectral envelopes between the encoder and the decoder. In addition, we relate the method to the linear predictive coding scheme and show its performance and robustness in a speech codec by objective and subjective evaluations. The codec based on this method, having bit-error robustness with only 1.5-ms algorithmic delay, can be useful at such situations as real-time speech communication with non-IP protocols.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
ICASSP2
2018 Optimal Golomb-Rice Code Extension for Lossless Coding of Low-Entropy Exponentially Distributed Sources
abstract
This paper presents an extension of GolombRice (GR) code for coding low-entropy sources, which the gap between their entropy and the conventional GR code length gets larger. We mention here the following four facts related to the proposed code, extended-domain GR (XDGR) code: it is represented by multiple code trees, based on the idea of almost instantaneous fixed-to-variable length codes, with its algorithm being a generalization of unary coding; its structure naturally contains run-length coding; the gap between the entropy and its average code length is theoretically guaranteed to be asymptotically negligible as the entropy of the exponentially distributed sources tends to zero; and its coding parameter, corresponding to the negative-domain Rice parameter of GR code, can be estimated from the input source-symbol sequence. Experimental evaluations are also presented supporting the theorems. The proposed XDGR code, having simple algorithm and high compression performance, is expected to be used for many coding applications, which deals with exponentially distributed sources at low bit rates.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
IEEE Trans. Inf. Theory2
2017 Shape parameter estimation for generalized-Gaussian-distributed frequency spectra of audio signals
abstract
We have devised a method for estimating, from a single frame of audio frequency spectra, a shape parameter of multivariate generalized Gaussian distribution which has variance represented by an all-pole model and no covariance. Based on powered all-pole spectrum estimation (PAPSE), which is an extension of linear prediction, the proposed method simultaneously estimates the shape parameter and the maximum-likelihood variance, allowing more accurate representation of the probability density functions of the spectra. This paper shows an integration of the estimation into an audio codec for an example of its application, which resulted in the enhancement of the objective and subjective reconstruction quality. Since this estimation method provides us with simple parameters which reflect some acoustic features of signals, the method may also be useful in other audio signal processing problems.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
ICASSP2
2015 Overview of the EVS codec architecture
abstract
The recently standardized 3GPP codec for Enhanced Voice Services (EVS) offers new features and improvements for low-delay real-time communication systems. Based on a novel, switched low-delay speech/audio codec, the EVS codec contains various tools for better compression efficiency and higher quality for clean/noisy speech, mixed content and music, including support for wideband, super-wideband and full-band content. The EVS codec operates in a broad range of bitrates, is highly robust against packet loss and provides an AMR-WB interoperable mode for compatibility with existing systems. This paper gives an overview of the underlying architecture as well as the novel technologies in the EVS codec and presents listening test results showing the performance of the new codec in terms of compression and speech/audio quality.
Martin Dietz, Markus Multrus, Vaclav Eksler, Vladimir Malenovsky, Erik Norvell, Harald Pobloth, Lei Miao 0004, Lasse Laaksonen, Adriana Vasilache, Yutaka Kamamoto, Kei Kikuiri, Stéphane Ragot, Julien Faure, Hiroyuki Ehara, Vivek Rajendran, Venkatraman Atti, Hosang Sung, Eunmi Oh, Changbao Zhu
ICASSP11
2015 Resolution Warped Spectral Representation for Low-Delay and Low-Bit-Rate Audio Coder
abstract
We have devised a high-quality frequency-domain audio coder based on the state-of-the-art monaural wide-band coder aiming at its use in low-delay and low-bit-rate conditions. The coder efficiently represents frequency spectral envelopes of the target signals with low computational complexity using optimally prepared non-negative sparse matrices. The experimental results reveal that this representation has positive effects on the objective and subjective quality of the coder resulting in the comparable quality to the same bit rate of 3GPP Extended Adaptive Multi-Rate WideBand ( AMR-WB+), a coder which permits more than four times longer delay compared with the proposed coder. Consequently, this coder is suitable for applications in mobile communications, which require low delay and low complexity.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Hirokazu Kameoka, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Optimal Coding of Generalized-Gaussian-Distributed Frequency Spectra for Low-Delay Audio Coder With Powered All-Pole Spectrum Estimation
abstract
We present an optimal coding scheme that parameterizes the maximum-likelihood estimate of variance for frequency spectra belonging to the generalized Gaussian distribution, the distribution covering the Laplacian and the Gaussian. By slightly modifying the all-pole model of the conventional linear prediction (LP), we can estimate the variance with the same method as in LP, which has low computational costs. Experimental results show that incorporating the coding scheme in a state-of-the-art wide-band audio coder enhances its objective and subjective quality in a low-bit-rate and low-delay situation by increasing the compression efficiency. Thus, this coding scheme will be useful in applications like mobile communications, which requires highly efficient compression.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Hirokazu Kameoka, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Simultaneous reconstruction of undersampled multichannel signals with a decayed and time-delayed common component
abstract
This paper presents a new signal model in distributed compressed sensing (DCS) and a reconstruction algorithm to reduce the number of samples required to accurately reconstruct signals obtained via sensors. While the conventional signal models in DCS assume that all signals include exactly the same common component, our model exploits the attenuation and time delay of the common component. With this extension, one can deal with signals originated from a source located in the field. To reconstruct such signals, we developed a new algorithm based on the alternative direction multiplier method. The algorithm efficiently reconstructs computer-generated signals using a reduced number of samples. In addition, we demonstrate the superior reconstruction quality with our method by using real electromyographic signals.
Yoshifumi Shiraki, Yutaka Kamamoto, Takehiro Moriya
ICASSP2
2010 Lossless Compression of Mapped Domain Linear Prediction Residual for ITU-T Recommendation G.711.0
abstract
Summary form only given. The ITU-T Rec. G.711 is widely used for narrowband telephony applications, including PSTN/GSTN and packet-based network applications such as VoIP, and has been used for many decades because of its proven voice quality, ubiquity, and utility. ITU has just established a lossless coding technology for G.711 encoded payloads, ITU-T Rec. G.711.0-Lossless compression of G.711 pulse code modulation. This paper introduces some coding technologies proposed and applied to the G.711.0 codec, such as Plus-Minus zero mapping for the mapped domain linear predictive coding and escaped-Huffman coding combined with adaptive recursive Rice coding for lossless compression of the prediction residual. The proposed and conventional algorithms were implemented and the coding performances of the codec were evaluated using large speech corpus in various conditions in terms of the computational complexity/compression performance trade-off. Since the distribution of the prediction residual signal sometimes does not follow the expected Laplacian PDF model of Rice coding, E-Huffman coding can improve the compression performance for relatively smaller Rice quotients of residual and Recursive Rice coding can improve the compression performance for relatively larger Rice quotients of residual. It is shown that the PM zero mapping improves the compression performance by 0.2% for ?-law input. The E-Huffman coding combined with adaptive recursive Rice coding improves the compression by 0.16% while the complexity was increased by 0.046 WMOPS for the encoder/decoder pair, averaged for all test conditions, compare to the conventional Rice coding scheme. The worst-case complexity was increased by 0.079 WMOPS. Average computational complexity is 1.071 WMOPS for the encoder/decoder pair and the worst-case complexity is 1.667 WMOPS in total.
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya
DCC2
2010 Low-Complexity PARCOR Coefficient Quantizer and Prediction Order Estimator for G.711.0 (Lossless Speech Coding)
abstract
This paper presents two low-complexity tools used for the new ITU-T recommendation G.711.0, which is the standard for lossless compression of G.711 (A-law/Mu-law logarithmic PCM) speech data. One is an algorithm for quantizing the PARCOR/reflection coefficients and the other is an estimation method for the optimal prediction order. Both tools are based on a criterion that minimizes the entropy of the prediction residual signals and can be implemented in a fixed-point low-complexity algorithm. G.711.0 with the developed practical tools will be widely used everywhere because it can losslessly reduce the data rate of G.711, the prevailing speech-coding technology.
Yutaka Kamamoto, Takehiro Moriya, Noboru Harada
DCC1
2010 Enhanced Lossless Coding Tools of LPC Residual for ITU-T G.711.0
abstract
Motivated by the rapid increase of VoIP services with G.711 for telephone speech, a new ITU-T recommendation, G.711.0 (frame-wise stateless lossless compression scheme for G.711 log PCM symbols), has been standardized. The standard scheme has several coding parts, each of which is adaptively selected depending on the characteristics of the input. Among them, the mapped domain prediction part is the one most frequently activated for normal speech signals. This part consists of linear prediction in the mapped domain and variable length coding of the prediction residual. It is useful for log-compressed/expanded signal, such as ITU-T G.711. This paper describes three newly devised enhancement tools for the coding of prediction residual signals: progressive order prediction, quantized prediction order, and adaptive and sub-frame base coding for separation parameters. The design criterion is the maximization of the averaged FoM (figure of merit) over frame lengths of 40, 80, 160, 240, and 320 samples. The first tool, progressive order prediction associated with the adaptive modification of the separation parameter for the first and second samples, enhances the compression ratio by 0.5 % with a negligible increase of the complexity. The second tool, quantized prediction order, improves the compression ratio by 0.2 % with even reduced complexity. The third tool, sub-frame base adaptive coding of separation parameters, gives a 0.2 % improvement in the compression ratio with comparable complexity. All three schemes are consistently and independently effective for improving the compression ratio, although the amount of improvement with each tool is small. At the same time, none of the tools have any significant impact on computational complexity. Therefore, all the devised tools improve the FoM and have been adopted in the mapped domain prediction part of the ITU-T G.711.0 standard.
Takehiro Moriya, Yutaka Kamamoto, Noboru Harada
DCC2
2010 Escaped-Huffman and adaptive recursive rice coding for lossless compression of the mapped domain linear prediction residual
abstract
ITU-T Recommendation G.711.0 has just been established. It defines a lossless and stateless compression for G.711 packet payloads (for both A-law and μ-law). This paper introduces some coding technologies proposed and applied to the G.711.0 codec, such as Plus-Minus zero mapping for the mapped domain linear predictive coding and escaped-Huffman coding combined with adaptive recursive Rice coding for lossless compression of the prediction residual. Performance test results for those coding tools are shown in comparison with the results for the conventional technology. The performance is measured based on the figure of merit (FoM), which is a function of the trade-off between compression performance and computational complexity. The proposed tools improve the compression performance by 0.16% in total while keeping the computational complexity of encoder/decoder pair low (about 1.0 WMOPS in average and 1.667 WMOPS in the worst-case).
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya
ICASSP2
2010 Emerging ITU-T standard G.711.0 - lossless compression of G.711 pulse code modulation
abstract
The ITU-T Recommendation G.711 is the benchmark standard for narrowband telephony. It has been successful for many decades because of its proven voice quality, ubiquity and utility. A new ITU-T recommendation, denoted G.711.0, has been recently established defining a lossless compression for G.711 packet payloads typically found in IP networks. This paper presents a brief overview of technologies employed within the G.711.0 standard and summarizes the compression and complexity results. It is shown that G.711.0 provides greater than 50% average compression in typical service provider environments while keeping low computational complexity for the encoder/decoder pair (1.0 WMOPS average, <;1.7 WMOPS worst case) and low memory footprint (about 5k octets RAM, 5.7k octets ROM, and 3.6k program memory measured in number of basic operators).
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya, Yusuke Hiwasaki, Michael A. Ramalho, Lorin Netsch, Jacek Stachurski, Lei Miao 0004, Hervé Taddei, Fengyan Qi
ICASSP2
2010 Low-complexity PARCOR coefficient quantizer and prediction order estimator for lossless speech coding
abstract
This paper describes two low-complexity tools used for the new ITU-T recommendation G.711.0, the lossless coding of G.711 (A-law/μ-law logarithmic PCM) speech data. One is an algorithm for quantizing the PARCOR/reflection coefficients and the other is an estimation method for the optimal prediction order. Both tools are based on a criterion that minimizes the entropy of the prediction residual signals and can be implemented in a fixed-point low-complexity algorithm. G.711.0 with the developed practical tools will be widely used everywhere because it can losslessly reduce the data rate of G.711, the prevailing speech-coding technology.
Yutaka Kamamoto, Takehiro Moriya, Noboru Harada
ICASSP1
2010 Enhanced lossless coding tools for prediction residual
abstract
Three elementary coding tools - a progressive order prediction tool, quantized order prediction tool, and adaptive and sub-frame base coding tool for separation parameters - have been devised to enhance the compression performance of the prediction residual. These are intended for the lossless coding of G.711 log PCM symbols used in packet-based network application such as VoIP. All tools are shown to be effective for reducing the average code length without any significant increase of computational complexity. As a result, all have been adopted in the mapped domain predictive coding part of the ITU-T G.711.0 standard.
Takehiro Moriya, Yutaka Kamamoto, Noboru Harada
ICASSP2
2008 Interchannel dependency analysis of biomedical signals for efficient lossless compression by MPEG-4 ALS
abstract
This paper describes a new search algorithm that quickly finds interchannel relationships between a coding channel and a reference channel in the multichannel coding tool of the MPEG- 4 Audio Lossless Coding (ALS) international standard. The algorithm has tree structure and can reduce data size with significantly smaller computation load than that of the conventional one. The devised method is based on a restricted greedy algorithm. It chooses the most efficient branch which does not make any loops in the existing path. The results of comprehensive evaluations show that this method maintains the compression performance (compression to around 1/3) and performs 1000 times as fast as the conventional method for the 512-channel magnetoencephalography signals. This algorithm enables practical lossless compression of biomedical data by the ALS, and at the same time, opens the way to a new multichannel analysis tool that may be used for purposes other than compression. The continual maintenance of this standard will make it possible to perfectly reconstruct encoded files even 100 years from now.
Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
ICASSP1
2008 Lossless compression of biomedical signals by MPEG-4 ALS with enhanced encoding tools
abstract
Enhanced encoding tools for MPEG-4 audio lossless coding (ALS) international standard were developed, with the goal of improving compression performance of time-series biomedical data. The multichannel coding (MCC) tool of this standard exploits interchannel redundancies to reduce the bit rate. To improve compression performance with the MCC, we have devised a multichannel linear prediction tool, which achieves around a 0.1% better compression ratio than that of the conventional method. We have also developed an interchannel dependency analysis tool, which performs about 1000 times faster than the conventional one. By combining these tools, biomedical signals are losslessly compressed to about 1/3 in a practical computational load. Compressed biomedical data will be decoded even 100 years from now, because the bitstream still remains compliant with the MPEG standard.
Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
MMSP1
2004 Complex spectrum circle centroid for microphone-array-based noisy speech recognition
abstract
We propose a novel principle based on Complex Spectrum Circle Centroid (CSCC) for restoring complex spectrum of the target signal from multiple microphone input signals in a noisy environment. If noise arrives at multiple microphones with different time delays relative to the target signal, the observed noisy signals lie on a circle in the complex spectrum plane from which the target signal is restored by finding the centroid of the circle. Unlike most of existing methods for noise reduction such as ICA, AMNOR and beamforming, this nonlinear operation is applicable to any type of noise including non-stationary, moving, signal-correlated, nonplanar, and spoken noises, without identifying the noise direction and training parameters.
Shigeki Sagayama, Okajima Takashi, Yutaka Kamamoto, Takuya Nishimoto
INTERSPEECH3