Takehiro Moriya

dblp:34/3176 · DBLP profile ↗
← Back
57ranked-venue papers
17as first author
4since 2021 · last 2025
0000-0003-4591-1273ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 14 first-author · 1 since 2021Artificial intelligence and machine learning · 11 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorTheory of computation · 4 · 2 since 2021Computer networks · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Stereo Downmix in 3GPP IVAS for EVS Compatibility
abstract
The 3GPP IVAS codec specifies an EVS-compatible stereo downmix as one of the key functionalities. This paper describes how this novel active downmix scheme has been devised to achieve high and stable quality from stereo input to EVS encoder/decoder with no additional algorithmic delay. An example of a network configuration for a multi-party conference using the proposed EVS-compatible downmix is also provided. This paper shows several fundamental schemes, of active processes, including adaptive weight and phase-compensated weight between two channels. Subjective listening test results show that the devised scheme’s quality is better than that of a passive downmix.
Takehiro Moriya, Stéphane Ragot, Arnaud Lefort, Alexandre Guérin, Noboru Harada, Ryosuke Sugiura, Yutaka Kamamoto
ICASSP1
2025 Optimal Construction of N-Bit-Delay Almost Instantaneous Fixed-to-Variable-Length Codes
abstract
This paper presents an optimal construction ofN-bit-delay almost instantaneous fixed-to-variable-length (AIFV) codes, the general form of binary codes we can make when finite bits of decoding delay are allowed. The presented method enables us to optimize lossless codes among a broader class of codes compared to the conventional FV and AIFV codes. The paper first discusses the problem of code construction, which contains some essential partial problems, and defines three classes of optimality to clarify how far we can solve the problems. The properties of the optimal codes are analyzed theoretically, showing the sufficient conditions for achieving the optimum. Then, we propose an algorithm for constructingN-bit-delay AIFV codes for given stationary memory-less sources. The optimality of the constructed codes is discussed both theoretically and empirically. They showed shorter expected code lengths whenN≥ 3 than the conventional AIFV-mand extended Huffman codes. Moreover, in the random numbers simulation, they performed higher compression efficiency than the 32-bit-precision range codes under reasonable conditions.
Ryosuke Sugiura, Masaaki Nishino, Norihito Yasuda, Yutaka Kamamoto, Takehiro Moriya
IEEE Trans. Inf. Theory5
2024 Device for controlling the phasic relationship between melodic sound and respiration and its effect on the change in respiration rate
abstract
Music is one of the most powerful mediators of awareness enhancement and respiration control.Although the timing (phase) control between music and respiration is important, few studies have focused on this.This study investigated the phasic relationship between melodic sound and respiration (PRSR) using a device that controls the PRSR.The interaction between music and respiration was realised in the phase by setting the template of the respiration pattern, called the 'target phase', which served as a guideline for coordinating the melodic phrase and respiration.In this experiment, we modelled six PRSR conditions by changing the start of the target phase.Participants were made to listen to cyclic melodies under the six PRSR conditions to examine the effect on their respiration interval.To modify the individual differences between participants, we realigned the PRSR conditions for each participant based on the error size between the respiration and target phases.We found a significant difference in the respiration intervals between the six realigned PRSR conditions, indicating the impact of the PRSR on the respiration rate and our effective control of PRSR.Our findings will help promote digital emotion regulation in the future, especially when using respiration as an input of feedback.
Takashi G. Sato, Yuuki Ooishi, Masahiro Fujino, Takehiro Moriya
Behav. Inf. Technol.4
2023 General Form of Almost Instantaneous Fixed-to-Variable-Length Codes
abstract
A general class of the almost instantaneous fixed-to-variable-length (AIFV) codes is proposed, which contains every possible binary code we can make when allowing finite bits of decoding delay. The contribution of the paper lies in the following. (i) Introducing$N$-bit-delay AIFV codes, constructed by multiple code trees with higher flexibility than the conventional AIFV codes. (ii) Proving that the proposed codes can represent any uniquely-encodable and uniquely-decodable variable-to-variable length codes. (iii) Showing how to express codes as multiple code trees with minimum decoding delay. (iv) Formulating the constraints of decodability as the comparison of intervals in the real number line. The theoretical results in this paper are expected to be useful for further study on AIFV codes.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
IEEE Trans. Inf. Theory3
2019 Phase Estimation Method Using Multiple Frames in Image-Sensor-Based Visible Light Communication
abstract
Visible light communication (VLC) is a wireless communication technology that employs visible light as a communication medium. Image-sensor-based VLC is suitable for multi-channel communication because multiple transmitters can be captured concurrently. In this paper, we propose a communication method for image-sensor-based VLC using multiple frames, which are intended to aid the simultaneous measurement of a multi-channel biological signal. The proposed method estimates the phase difference between the transmitter and the receiver using multiple frames shot at different timings compared to LED blinking, and communication is established through phase shift keying. Moreover, the proposed method is resistant to phase differences between the transmitter and the receiver, which is unavoidable in asynchronous communication. Through experiments, we confirmed that communication is possible using our proposed method.
Shohei Maruyama, Tomohiro Yendo, Yoshifumi Shiraki, Takashi G. Sato, Takehiro Moriya
CCNC5
2019 Detecting Attention Shift from Neural Response Based on Beat-frequency-modulated Musical Excerpts
abstract
This paper presents a new approach for detecting attention from auditory steady-state responses (ASSR) by using musical excerpts. The feature extraction process for electroencephalogram (EEG) signal is combined with a support vector machine as a binary discriminator. A novel modulation that emphasizes the beat timing of the excerpts has enhances the EEG response. Thanks to the beat-locked epoch extraction and additional information that signals the locked excerpt, the estimation errors are less than 8% using only ten seconds of data. Contrary to our expectations, a waveform-averaging method outperforms a harmonic filter bank and a bin-subtraction methods in the frequency domain for feature extraction. Overall, attention estimation from EEGs using musical excerpts as stimuli has been successfully achieved, which represents significant progress towards the development of a mass-EEG measurement system.
Takashi G. Sato, Yoshifumi Shiraki, Takehiro Moriya
ICASSP3
2019 Shape Control of Discrete Generalized Gaussian Distributions for Frequency-Domain Audio Coding
abstract
Entropy coding, which is an essential part of audio compression, is always required to manage the tradeoffs between compression efficiency and computational complexity, and the strategy to achieve them highly depends on the distributions of inputs. In this paper, we present a method of controlling them for enhancing the compression efficiency of Golomb-Rice (GR) encoding, one of the simplest entropy coding methods optimal for Laplacian distributions. We will show that the proposed invertible and low-complexity mapping of integers enables the GR encoding to assign nearly the optimal code length for a wider range of distributions, generalized Gaussian distributions, maintaining low computational cost. A simulation by random numbers reveals that the proposed coder based on this scheme works about 6 times faster than the state-of-the-art arithmetic coder for Gaussian-distributed integers maintaining the increase in relative redundancy around 2.6%, which is much lower than that of a conventional GR coder. Additionally, an application to a practical speech and audio coding scheme is presented, and an objective evaluation for real speech and audio signals confirms the advantages of the proposed method in compression. The method is expected to widen the capability of low-complexity entropy coding, providing us with more flexible codec designs.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.3
2018 Integer Nesting/Splitting for Golomb-Rice Coding of Generalized Gaussian Sources
abstract
This paper presents a qualitative approach of combining Golomb-Rice (GR) code with algebraic bijective mappings which losslessly convert between arbitrary positive integers of different dimension and shape the distribution of generalized Gaussian sources. The mappings, integer nesting and splitting, enables GR encoding, with a little additional computation, to compress more efficiently sources based on wider classes of distributions than Laplacian. Simulations showed, especially for some Gaussian sources, almost optimal average code length can be achievable by performing integer nesting before GR encoding the integers. This scheme will be useful for applications dealing with various types of sources and requiring low computational costs.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
DCC3
2018 Spectral-Envelope-Based Least Significant Bit Management for Low-Delay Bit-Error-Robust Speech Coding
abstract
We have devised a method for bit assignment of quantized frequency spectra aiming at its use in low-delay bit-error-robust speech compression. The proposed method, least significant bit management (LSBM), controls the least significant bits of the spectra based on their envelope to make them represented by fixed bit rates, which guarantees the range of the damage caused by the bit error which makes the mismatch of the spectral envelopes between the encoder and the decoder. In addition, we relate the method to the linear predictive coding scheme and show its performance and robustness in a speech codec by objective and subjective evaluations. The codec based on this method, having bit-error robustness with only 1.5-ms algorithmic delay, can be useful at such situations as real-time speech communication with non-IP protocols.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
ICASSP3
2018 Optimal Golomb-Rice Code Extension for Lossless Coding of Low-Entropy Exponentially Distributed Sources
abstract
This paper presents an extension of GolombRice (GR) code for coding low-entropy sources, which the gap between their entropy and the conventional GR code length gets larger. We mention here the following four facts related to the proposed code, extended-domain GR (XDGR) code: it is represented by multiple code trees, based on the idea of almost instantaneous fixed-to-variable length codes, with its algorithm being a generalization of unary coding; its structure naturally contains run-length coding; the gap between the entropy and its average code length is theoretically guaranteed to be asymptotically negligible as the entropy of the exponentially distributed sources tends to zero; and its coding parameter, corresponding to the negative-domain Rice parameter of GR code, can be estimated from the input source-symbol sequence. Experimental evaluations are also presented supporting the theorems. The proposed XDGR code, having simple algorithm and high compression performance, is expected to be used for many coding applications, which deals with exponentially distributed sources at low bit rates.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
IEEE Trans. Inf. Theory4
2017 Shape parameter estimation for generalized-Gaussian-distributed frequency spectra of audio signals
abstract
We have devised a method for estimating, from a single frame of audio frequency spectra, a shape parameter of multivariate generalized Gaussian distribution which has variance represented by an all-pole model and no covariance. Based on powered all-pole spectrum estimation (PAPSE), which is an extension of linear prediction, the proposed method simultaneously estimates the shape parameter and the maximum-likelihood variance, allowing more accurate representation of the probability density functions of the spectra. This paper shows an integration of the estimation into an audio codec for an example of its application, which resulted in the enhancement of the objective and subjective reconstruction quality. Since this estimation method provides us with simple parameters which reflect some acoustic features of signals, the method may also be useful in other audio signal processing problems.
Ryosuke Sugiura, Yutaka Kamamoto, Takehiro Moriya
ICASSP3
2017 Audience excitement reflected in respiratory phase synchronization
abstract
This paper investigates respiration activity data obtained simultaneously from more than ten persons while they were attending a live concert. Physiological activity in a group setting is attracting increased attention since it was suggested that it differs from data obtained individually with recorded stimuli. To overcome the difficulty in obtaining physiological data in a real situation, we developed a respiratory measurement system based on optical camera communication. We used the system to simultaneously measure respiration data from 15 participants who were in the audience of a live concert. Using the data, we showed that the respiratory phase distribution was distorted during some of the program suggesting a synchronization effect. We also found a program that was exciting for the participants and a program where the synchronization effects were almost identical, which suggests that a respiratory phase index can possibly be used to evaluate the quality of concerts.
Takashi G. Sato, Yoshifumi Shiraki, Takehiro Moriya
SMC3
2015 Low delay LPC and MDCT-based audio coding in the EVS codec
abstract
Speech coders operating in time domain can be extended with a frequency domain mode to improve encoding of music, even though this is challenging at low delay. In such a scenario, the short analysis window limits the benefit of the transform coder, while a delayless switch between the two coders constrains the system further. The paper presents an LPC and MDCT-based audio coder part of the new 3GPP codec for Enhanced Voice Services, which aims to solve the issues. Several advanced coding tools are introduced to alleviate the constraints: transient handling is improved, harmonic structures are better preserved, and the modeling of the zero-quantized frequencies is enhanced. Test results show that the obtained low-delay switched coder brings a clear improvement over a speech coder and is competitive even in comparison to audio coders with higher delay.
Guillaume Fuchs, Christian R. Helmrich, Goran Markovic, Matthias Neusinger, Emmanuel Ravelli, Takehiro Moriya
ICASSP6
2015 Resolution Warped Spectral Representation for Low-Delay and Low-Bit-Rate Audio Coder
abstract
We have devised a high-quality frequency-domain audio coder based on the state-of-the-art monaural wide-band coder aiming at its use in low-delay and low-bit-rate conditions. The coder efficiently represents frequency spectral envelopes of the target signals with low computational complexity using optimally prepared non-negative sparse matrices. The experimental results reveal that this representation has positive effects on the objective and subjective quality of the coder resulting in the comparable quality to the same bit rate of 3GPP Extended Adaptive Multi-Rate WideBand ( AMR-WB+), a coder which permits more than four times longer delay compared with the proposed coder. Consequently, this coder is suitable for applications in mobile communications, which require low delay and low complexity.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Hirokazu Kameoka, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.5
2015 Optimal Coding of Generalized-Gaussian-Distributed Frequency Spectra for Low-Delay Audio Coder With Powered All-Pole Spectrum Estimation
abstract
We present an optimal coding scheme that parameterizes the maximum-likelihood estimate of variance for frequency spectra belonging to the generalized Gaussian distribution, the distribution covering the Laplacian and the Gaussian. By slightly modifying the all-pole model of the conventional linear prediction (LP), we can estimate the variance with the same method as in LP, which has low computational costs. Experimental results show that incorporating the coding scheme in a state-of-the-art wide-band audio coder enhances its objective and subjective quality in a low-bit-rate and low-delay situation by increasing the compression efficiency. Thus, this coding scheme will be useful in applications like mobile communications, which requires highly efficient compression.
Ryosuke Sugiura, Yutaka Kamamoto, Noboru Harada, Hirokazu Kameoka, Takehiro Moriya
IEEE ACM Trans. Audio Speech Lang. Process.5
2013 Simultaneous reconstruction of undersampled multichannel signals with a decayed and time-delayed common component
abstract
This paper presents a new signal model in distributed compressed sensing (DCS) and a reconstruction algorithm to reduce the number of samples required to accurately reconstruct signals obtained via sensors. While the conventional signal models in DCS assume that all signals include exactly the same common component, our model exploits the attenuation and time delay of the common component. With this extension, one can deal with signals originated from a source located in the field. To reconstruct such signals, we developed a new algorithm based on the alternative direction multiplier method. The algorithm efficiently reconstructs computer-generated signals using a reduced number of samples. In addition, we demonstrate the superior reconstruction quality with our method by using real electromyographic signals.
Yoshifumi Shiraki, Yutaka Kamamoto, Takehiro Moriya
ICASSP3
2010 Lossless Compression of Mapped Domain Linear Prediction Residual for ITU-T Recommendation G.711.0
abstract
Summary form only given. The ITU-T Rec. G.711 is widely used for narrowband telephony applications, including PSTN/GSTN and packet-based network applications such as VoIP, and has been used for many decades because of its proven voice quality, ubiquity, and utility. ITU has just established a lossless coding technology for G.711 encoded payloads, ITU-T Rec. G.711.0-Lossless compression of G.711 pulse code modulation. This paper introduces some coding technologies proposed and applied to the G.711.0 codec, such as Plus-Minus zero mapping for the mapped domain linear predictive coding and escaped-Huffman coding combined with adaptive recursive Rice coding for lossless compression of the prediction residual. The proposed and conventional algorithms were implemented and the coding performances of the codec were evaluated using large speech corpus in various conditions in terms of the computational complexity/compression performance trade-off. Since the distribution of the prediction residual signal sometimes does not follow the expected Laplacian PDF model of Rice coding, E-Huffman coding can improve the compression performance for relatively smaller Rice quotients of residual and Recursive Rice coding can improve the compression performance for relatively larger Rice quotients of residual. It is shown that the PM zero mapping improves the compression performance by 0.2% for ?-law input. The E-Huffman coding combined with adaptive recursive Rice coding improves the compression by 0.16% while the complexity was increased by 0.046 WMOPS for the encoder/decoder pair, averaged for all test conditions, compare to the conventional Rice coding scheme. The worst-case complexity was increased by 0.079 WMOPS. Average computational complexity is 1.071 WMOPS for the encoder/decoder pair and the worst-case complexity is 1.667 WMOPS in total.
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya
DCC3
2010 Low-Complexity PARCOR Coefficient Quantizer and Prediction Order Estimator for G.711.0 (Lossless Speech Coding)
abstract
This paper presents two low-complexity tools used for the new ITU-T recommendation G.711.0, which is the standard for lossless compression of G.711 (A-law/Mu-law logarithmic PCM) speech data. One is an algorithm for quantizing the PARCOR/reflection coefficients and the other is an estimation method for the optimal prediction order. Both tools are based on a criterion that minimizes the entropy of the prediction residual signals and can be implemented in a fixed-point low-complexity algorithm. G.711.0 with the developed practical tools will be widely used everywhere because it can losslessly reduce the data rate of G.711, the prevailing speech-coding technology.
Yutaka Kamamoto, Takehiro Moriya, Noboru Harada
DCC2
2010 Enhanced Lossless Coding Tools of LPC Residual for ITU-T G.711.0
abstract
Motivated by the rapid increase of VoIP services with G.711 for telephone speech, a new ITU-T recommendation, G.711.0 (frame-wise stateless lossless compression scheme for G.711 log PCM symbols), has been standardized. The standard scheme has several coding parts, each of which is adaptively selected depending on the characteristics of the input. Among them, the mapped domain prediction part is the one most frequently activated for normal speech signals. This part consists of linear prediction in the mapped domain and variable length coding of the prediction residual. It is useful for log-compressed/expanded signal, such as ITU-T G.711. This paper describes three newly devised enhancement tools for the coding of prediction residual signals: progressive order prediction, quantized prediction order, and adaptive and sub-frame base coding for separation parameters. The design criterion is the maximization of the averaged FoM (figure of merit) over frame lengths of 40, 80, 160, 240, and 320 samples. The first tool, progressive order prediction associated with the adaptive modification of the separation parameter for the first and second samples, enhances the compression ratio by 0.5 % with a negligible increase of the complexity. The second tool, quantized prediction order, improves the compression ratio by 0.2 % with even reduced complexity. The third tool, sub-frame base adaptive coding of separation parameters, gives a 0.2 % improvement in the compression ratio with comparable complexity. All three schemes are consistently and independently effective for improving the compression ratio, although the amount of improvement with each tool is small. At the same time, none of the tools have any significant impact on computational complexity. Therefore, all the devised tools improve the FoM and have been adopted in the mapped domain prediction part of the ITU-T G.711.0 standard.
Takehiro Moriya, Yutaka Kamamoto, Noboru Harada
DCC1
2010 Escaped-Huffman and adaptive recursive rice coding for lossless compression of the mapped domain linear prediction residual
abstract
ITU-T Recommendation G.711.0 has just been established. It defines a lossless and stateless compression for G.711 packet payloads (for both A-law and μ-law). This paper introduces some coding technologies proposed and applied to the G.711.0 codec, such as Plus-Minus zero mapping for the mapped domain linear predictive coding and escaped-Huffman coding combined with adaptive recursive Rice coding for lossless compression of the prediction residual. Performance test results for those coding tools are shown in comparison with the results for the conventional technology. The performance is measured based on the figure of merit (FoM), which is a function of the trade-off between compression performance and computational complexity. The proposed tools improve the compression performance by 0.16% in total while keeping the computational complexity of encoder/decoder pair low (about 1.0 WMOPS in average and 1.667 WMOPS in the worst-case).
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya
ICASSP3
2010 Emerging ITU-T standard G.711.0 - lossless compression of G.711 pulse code modulation
abstract
The ITU-T Recommendation G.711 is the benchmark standard for narrowband telephony. It has been successful for many decades because of its proven voice quality, ubiquity and utility. A new ITU-T recommendation, denoted G.711.0, has been recently established defining a lossless compression for G.711 packet payloads typically found in IP networks. This paper presents a brief overview of technologies employed within the G.711.0 standard and summarizes the compression and complexity results. It is shown that G.711.0 provides greater than 50% average compression in typical service provider environments while keeping low computational complexity for the encoder/decoder pair (1.0 WMOPS average, <;1.7 WMOPS worst case) and low memory footprint (about 5k octets RAM, 5.7k octets ROM, and 3.6k program memory measured in number of basic operators).
Noboru Harada, Yutaka Kamamoto, Takehiro Moriya, Yusuke Hiwasaki, Michael A. Ramalho, Lorin Netsch, Jacek Stachurski, Lei Miao 0004, Hervé Taddei, Fengyan Qi
ICASSP3
2010 Low-complexity PARCOR coefficient quantizer and prediction order estimator for lossless speech coding
abstract
This paper describes two low-complexity tools used for the new ITU-T recommendation G.711.0, the lossless coding of G.711 (A-law/μ-law logarithmic PCM) speech data. One is an algorithm for quantizing the PARCOR/reflection coefficients and the other is an estimation method for the optimal prediction order. Both tools are based on a criterion that minimizes the entropy of the prediction residual signals and can be implemented in a fixed-point low-complexity algorithm. G.711.0 with the developed practical tools will be widely used everywhere because it can losslessly reduce the data rate of G.711, the prevailing speech-coding technology.
Yutaka Kamamoto, Takehiro Moriya, Noboru Harada
ICASSP2
2010 Enhanced lossless coding tools for prediction residual
abstract
Three elementary coding tools - a progressive order prediction tool, quantized order prediction tool, and adaptive and sub-frame base coding tool for separation parameters - have been devised to enhance the compression performance of the prediction residual. These are intended for the lossless coding of G.711 log PCM symbols used in packet-based network application such as VoIP. All tools are shown to be effective for reducing the average code length without any significant increase of computational complexity. As a result, all have been adopted in the mapped domain predictive coding part of the ITU-T G.711.0 standard.
Takehiro Moriya, Yutaka Kamamoto, Noboru Harada
ICASSP1
2008 Interchannel dependency analysis of biomedical signals for efficient lossless compression by MPEG-4 ALS
abstract
This paper describes a new search algorithm that quickly finds interchannel relationships between a coding channel and a reference channel in the multichannel coding tool of the MPEG- 4 Audio Lossless Coding (ALS) international standard. The algorithm has tree structure and can reduce data size with significantly smaller computation load than that of the conventional one. The devised method is based on a restricted greedy algorithm. It chooses the most efficient branch which does not make any loops in the existing path. The results of comprehensive evaluations show that this method maintains the compression performance (compression to around 1/3) and performs 1000 times as fast as the conventional method for the 512-channel magnetoencephalography signals. This algorithm enables practical lossless compression of biomedical data by the ALS, and at the same time, opens the way to a new multichannel analysis tool that may be used for purposes other than compression. The continual maintenance of this standard will make it possible to perfectly reconstruct encoded files even 100 years from now.
Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
ICASSP3
2008 Lossless compression of biomedical signals by MPEG-4 ALS with enhanced encoding tools
abstract
Enhanced encoding tools for MPEG-4 audio lossless coding (ALS) international standard were developed, with the goal of improving compression performance of time-series biomedical data. The multichannel coding (MCC) tool of this standard exploits interchannel redundancies to reduce the bit rate. To improve compression performance with the MCC, we have devised a multichannel linear prediction tool, which achieves around a 0.1% better compression ratio than that of the conventional method. We have also developed an interchannel dependency analysis tool, which performs about 1000 times faster than the conventional one. By combining these tools, biomedical signals are losslessly compressed to about 1/3 in a practical computational load. Compressed biomedical data will be decoded even 100 years from now, because the bitstream still remains compliant with the MPEG standard.
Yutaka Kamamoto, Noboru Harada, Takehiro Moriya
MMSP3
2004 Extended linear prediction tools for lossless audio coding
abstract
Two extension tools for enhancing the compression performance of prediction-based lossless audio coding are proposed. One is progressive-order prediction of the starting samples at the random access points, where the information of previous samples is not available. The first sample is coded as is, the second is predicted by first-order prediction, the third is predicted by second-order prediction, and so on. This can be efficiently carried out with PAR-COR (PARtial autoCORrelation) coefficients. The second tool is interchannel joint coding. Both predictive coefficients and prediction error signals are efficiently coded by interchannel differential or three-tap adaptive prediction. These new prediction tools lead to a steady reduction in bit rate when random access is activated and the interchannel correlation is strong.
Takehiro Moriya, Dai Yang, Tilman Liebchen
ICASSP (3)1
2004 A lossless audio compression scheme with random access property
abstract
We propose an efficient lossless coding algorithm that not only handles both PCM format data and IEEE floating-point format data, but also provides end users with a random access property. In the worst-case scenario, where the proposed algorithm was applied to artificially generated full 32 bit floating-point sound files with 48 kHz or 96 kHz sampling frequencies, an average compression rate of more than 1.5 and 1.7, respectively, was still achieved, which is much better than the average compression rate of less than 1.1 achieved by the general purpose lossless coding algorithm, gzip. Moreover, input sound files with samples' magnitudes out-of-range can also be perfectly reconstructed by our algorithm.
Dai Yang, Takehiro Moriya, Tilman Liebchen
ICASSP (3)2
2003 Hierarchical lossless audio coding in terms of sampling rate and amplitude resolution
abstract
This paper proposes a lossless audio coding scheme with hierarchical scalability in terms of sampling rate and amplitude resolution. A single bit stream contains hierarchical information that can generate waveforms ranging from 96 kHz with 24-bit amplitude resolution through lower sampling/resolution lossless waveforms to a highly compressed lossy one created using an MPEG-4 audio coder. This bit stream structure enables dynamic rate control and hierarchical multicasting based on a simple priority control of the IP packets. These functions will be useful for high-quality archiving and broadband streaming for various types of networks and terminal equipment.
Takehiro Moriya, Akio Jin, Takeshi Mori, Kazunaga Ikeda, Takao Kaneko
ICASSP (5)1
2002 Lossless scalable audio coder and quality enhancement
abstract
This paper proposes a lossless scalable audio coding scheme and quality enhancement processing at the decoder to compensate for some missing scalable units of information. The bit rate scalability is achieved by combining high-compression coding, such as MPEG-4, and horizontal bit slicing of the PCM-coded error signal between the original waveform and the locally reconstructed MPEG-4 signal. The horizontally sliced stream may be transported through an IP network with priority. Even if some units are missing at the decoder, reasonable quality waveform can be reconstructed by means of post-processing using some side information. This scheme enables graceful degradation by supporting lossless, near lossless, and high-compression coding within a single scalable framework, and is useful for narrowband to broadband audio streaming.
Takehiro Moriya, Akio Jin, Takeshi Mori, Kazunaga Ikeda, Takao Kaneko
ICASSP1
2001 Fast encoding algorithms for MPEG-4 TwinVQ audio tool
abstract
The ISO/IEC MPEG-4 audio standard includes the TwinVQ encoding tool. This tool is suitable for low-bit-rate general audio coding, but a drawback is the computational complexity of the encoder. To develop a faster TwinVQ encoder, new fast vector quantization algorithms - area localized pre-selection and hit zone masking - are introduced. These algorithms exploit the pre- and main-selection procedure scheme of the conjugate structure vector quantization which is used in the TwinVQ. The improvement is evaluated by measuring the encoding speed and the sound quality of reproduction.
Naoki Iwakami, Takehiro Moriya, Akio Jin, Takeshi Mori, Kazuaki Chikira
ICASSP2
2000 A design of lossy and lossless scalable audio coding
abstract
This paper proposes a lossless audio coding scheme making use of high-compression lossy coding such as the ISO/IEC MPEG-4 audio standard. The encoding process consists of lossy compression, bit slice lossless data conversion, and lossless coding. The lossy compression significantly reduces the amplitude of the error signal between the input signal and the reconstructed signal. In addition, bit slice data conversion makes the lossless coding more efficient. As a result of preliminary experiment, a total file size can be reduced to 70-50% of the original size, which is significantly more efficient than conventional universal lossless coding schemes such as "gzip". The resultant compressed data is useful for various applications, since it has bit rate scalability and can be used as a highly compressed signal as well as a lossless signal.
Takehiro Moriya, Naoki Iwakami, Akio Jin, Takeshi Mori
ICASSP1
2000 A design of error robust scalable coder based on MPEG-4/Audio
abstract
This paper proposes a combined design for an error robust and scalable audio coder based on TwinVQ (transform domain weighted interleave vector quantization) object type and EP (error protection) tools defined in MPEG-4/Audio. By combining the inherent error-robustness of TwinVQ, the un-equal error-protection scheme, and the hierarchical scalable coding structure, the distortion is confirmed to be very minor, even in high-error channel conditions at bit rates of 16-32 kbit/s. These structures offer a flexible coding framework which can be used for various types of network/storage environment and applications.
Takehiro Moriya, Takeshi Mori, Naoki Iwakami, Akio Jin
ISCAS1
1999 Scalable audio coder based on quantizer units of MDCT coefficients
abstract
A scalable codec has been constructed by using transform coding and the basic modules for scalable encoder and decoder. It allows users to choose a variety of scalable configurations in the frequency domain. The basic module is a quantizer that can quantize MDCT (modified DCT) coefficients transformed from a variety of frequency regions. This module mainly works at bit rates of more than 8 kbit/s. We can also change the target frequency regions of the basic module's input-output signals in each transform frame; i.e., we can change the scalable structure according to the nature of the input signals. In the scalable codec described here, the input-output signals are monaural and the sampling frequency is 24 kHz. The total bit rate of this scalable codec is more than 8 kbit/s. Subjective quality evaluation tests, mainly for musical sound sources, showed that it's sound quality is better than that of an MPEG-2 layer 3 codec at 8, 16, and 24 kbit/s when our scalable codec is constructed of 8-kbit/s basic modules. In combination with AAC (advanced audio coding), our scalable codec will be chosen as an international standard in ISO/IEC-MPEG-4/Audio.
Akio Jin, Takehiro Moriya, Takeshi Norimatsu, Mineo Tsushima, Tomokazu Ishikawa
ICASSP2
1998 Design and description of CS-ACELP: a toll quality 8 kb/s speech coder
abstract
This paper describes the 8 kb/s speech coding algorithm G.729 which has been standardized by ITU-T. The algorithm is based on a conjugate-structure algebraic CELP (CS-ACELP) coding technique and uses 10 ms speech frames. The codec delivers toll-quality speech (equivalent to 32 kb/s ADPCM) for most operating conditions. This paper describes the coder structure in detail and discusses the reasons behind certain design choices. A 16-b fixed-point version has been developed as part of Recommendation G.729 and a summary of the subjective test results based on a real-time implementation of this version are presented.
Redwan Salami, Claude Laflamme, Jean-Pierre Adoul, Akitoshi Kataoka, Shinji Hayashi, Takehiro Moriya, Claude Lamblin, Dominique Massaloux, Stéphane Proust, Peter Kroon, Yair Shoham
IEEE Trans. Speech Audio Process.6
1997 A design of transform coder for both speech and audio signals at 1 bit/sample
abstract
This paper proposes a speech and audio coder which operates at 1 bit/sample, namely an 8 kbit/s coder for 8 kHz sampling or a 16 kbit/s coder for 16 kHz sampling. The basic structure is inherited from a Twin VQ (transform domain weighted interleave vector quantization) high-quality audio coding scheme. A periodical component extraction scheme is newly added to the quantization of the MDCT coefficients. This scheme is found to be effective for reducing distortion and improving the robustness against channel errors. The qualities for music signals at 8 kbit/s are better than those of G.729 at the same bit rates, while they are worse for clean speech. The qualities at 16 kbit/s are comparable to or better than those of G.722 at 48 kbit/s.
Takehiro Moriya, Naoki Iwakami, Akio Jin, Kazunaga Ikeda, Satoshi Miki
ICASSP1
1996 Extension and complexity reduction of TwinVQ audio coder
abstract
This paper proposes two novel techniques for twinVQ (transform domain weighted interleave VQ) high-quality audio coding scheme for rates lower than 64 kbit/s. One is an extension of the weighted interleave technique to the time and input channel domains as well as the frequency domain. The other is an efficient representation scheme of the spectral envelope by means of a interpolated square root LPC (linear predictive coding) spectrum.
Takehiro Moriya, Naoki Iwakami, Kazunaga Ikeda, Satoshi Miki
ICASSP1
1996 Dual-Pulse CS-CELP: a toll-quality low-complexity speech coder at 7.8 kbit/s
abstract
Low-cost highly-efficient speech coding is important for personal multimedia communications. In this paper we propose a new low-complexity speech coding method called "Dual-Pulse CS-CELP (DP-CS-CELP)" at 7.8 kbit/s. This method is based on ITU-T G.729. To reduce the complexity, we applied a new excitation model to the random code vectors and simplified the LPC coding, adaptive codebook search, perceptual weighting, and the other structures. The number of operations for this encoder is 3.80 MOPS, which achieves real-time speech coding and decoding on personal computers. Although the efficiency of this coder is a little lower than that of G.729, the MOS listening test showed that the subjective quality was equivalent to or a little better than that of G.726 (32-kbit/s ADPCM).
Hitoshi Ohmuro, Jotaro Ikedo, Takehiro Moriya, Akitoshi Kataoka, Shinji Hayashi, Kazunori Mano
ICASSP3
1996 An 8-kb/s conjugate structure CELP (CS-CELP) speech coder
abstract
This paper describes a high-quality 8-kb/s speech coder called conjugate structure code-excited linear prediction (CS-CELP) with a 10-ms frame length. To provide a short delay and high quality under both error-free and channel error conditions, it uses three new schemes: line spectrum pair (LSP) quantization using interframe prediction, preselection in the codebook search, and gain vector quantization (VQ) with backward prediction. The LSP parameters are quantized by using multistage VQ with moving-average (MA) prediction. This scheme can operate efficiently with various frequency responses of speech. The preselection of the codebook reduces the computational complexity and improves the robustness to channel errors. The gain VQ with backward prediction can provide a high quality and robustness without transmission of input speech power information. A conjugate structure for both random codebook and gain codebook is introduced to improve the ability to handle random bit errors and to reduce codebook storage memory requirements. Subjective testing indicates that the quality of this coder is equivalent to that of 32-kb/s adaptive differential pulse code modulation (ADPCM) under error-free conditions. Testing has further demonstrated that the coder is robust against random bit errors.
Akitoshi Kataoka, Takehiro Moriya, Shinji Hayashi
IEEE Trans. Speech Audio Process.2
1995 High-quality audio-coding at less than 64 kbit/s by using transform-domain weighted interleave vector quantization (TwinVQ)
abstract
A new audio-coding method is proposed. This method is called transform-domain weighted interleave vector quantization (TwinVQ) and achieves high-quality reproduction at less than 64 kbit/s. The method is a transform coding using modified discrete cosine transform (MDCT). There are three novel techniques in this method: flattening of the MDCT coefficients by the spectrum of linear predictive coding (LPC) coefficients; interframe backward prediction for flattening the MDCT coefficients; and weighted interleave vector quantization. Subjective evaluation tests showed that the quality of the reproduction of TwinVQ exceeded that of an MPEG Layer II coder at the same bitrate.
Naoki Iwakami, Takehiro Moriya, Satoshi Miki
ICASSP2
1995 Improved CS-CELP speech coding in a noisy environment using a trained sparse conjugate codebook
abstract
A high-quality 8-kbit/s speech coder based on conjugate structure CELP (CS-CELP) is proposed that uses a trained sparse conjugate codebook. The trained sparse conjugate codebook improves speech quality for noisy speech. This codebook consists of two sub-codebooks and each sub-codebook consists of a random component and a trained component. Each component has excitation vectors consisting of a few pulses. In the random component, pulse position and amplitude are determined randomly. The trained component is determined by training. Subjective tests (differential mean opinion score, DMOS and mean opinion score, MOS) indicated that this codebook improves speech quality compared with the conventional trained codebook for noisy speech. The MOS showed that the quality of improved CS-CELP is equivalent to that of the 32-kbit/s ADPCM for clean speech.
Akitoshi Kataoka, Sachiko Hosaka, Jotaro Ikedo, Takehiro Moriya, Shinji Hayashi
ICASSP4
1995 Wideband CELP coder at 16-kbit/s with 10-ms frame
Shigeaki Sasaki, Akitoshi Kataoka, Takehiro Moriya
EUROSPEECH3
1995 Design of a Pitch Synchronous Innovation CELP Coder for Mobile Communications
abstract
This paper describes the design of a speech coder called pitch synchronous innovation CELP (PSI-CELP) for low hit-rate mobile communications. PSI-CELP is based on CELP, but has more adaptive excitation structures. In voiced frames, instead of conventional random excitation vectors, PSI-CELP converts even the random excitation vectors to have pitch periodicity by repeating stored random vectors as well as by using an adaptive codebook, in silent, unvoiced, and transient frames, the coder stops using the adaptive codebook and switches to fixed random codebooks. The PSI-CELP coder also implements novel structures and techniques: an FIR-type perceptual weighting filter using unquantized LPC parameters, a random codebook with a conjugate structure trained to be robust against channel errors, codebook search with delayed decision, a gain quantization with sloped amplitude, and a moving average prediction coding of LSP parameters, Our speech coder is implemented by DSP chips. Its coded speech quality at 3.6 kb/s with 2.0 kb/s redundancy is comparable to that of the Japanese full-rate VSELP coder at 6.7 kb/s with 4.5 kb/s redundancy. The basic structure of this PSI-CELP coder has been chosen as the Japanese half-rate speech codec for digital cellular telecommunications.>
Kazunori Mano, Takehiro Moriya, Satoshi Miki, Hitoshi Ohmuro, Kazunaga Ikeda, Jotaro Ikedo
IEEE J. Sel. Areas Commun.2
1994 Implementation and performance of an 8-kbit/s conjugate structure CELP speech coder
abstract
This paper presents a high-quality 8-kbit/s speech coder (conjugate structure CELP: CS-CELP) that is a candidate for standardization by the ITU-T (formerly CCITT). To achieve high-quality for two types of speech (IRS and non-IRS (flat) speech) and real-time implementation, CS-CELP has been revised by two novel schemes. To handle two types of speech, the LSP parameters are quantized by multistage VQ with fourth-order interframe MA prediction. This scheme has little spectrum distortion, even if the two types of speech have many variations of the LSP parameters. The computational complexity of the implementation is reduced for adaptive and fixed-shape codebooks without degrading the speech quality. Multistage selection is adopted in the adaptive codebook; this selection uses a truncated impulse response. Improved pre-selection is proposed in the fixed-shape codebook. Subjective testing indicates that the quality of CS-CELP is equivalent to that of the 32-kbit/s ADPCM under error-free conditions for IRS and non-IRS speech. It also operates in real time using fixed-point DSP chips.>
Akitoshi Kataoka, Takehiro Moriya, Shinji Hayashi
ICASSP (2)2
1994 A pitch synchronous innovation CELP (PSI-CELP) coder for 2-4 kbit/s
abstract
This paper proposes high-quality and low bit-rate (3.6 and 2.4 kbit/s) coders using a pitch synchronous innovation CELP (PSI-CELP) method or a phase adaptive PSI-CELP. PSI-CELP, which is used as the excitation structure of the half-rate codec for the standard of Japanese digital mobile telephony, is based on CELP but adds pitch synchronous innovation, which means that even random codevectors are adaptively converted to have pitch periodicity for voiced frames. Phase adaptive PSI-CELP makes not only the periodicity, like in PSI-CELP, but also the phase of random codevectors equal to those of an adaptive codevector. The subjective qualities of the 3.6- and 2.4-kbit/s coders exceed those of the 6.7-kbit/S VSELP coder, which is the full-rate codec for the standard of Japanese digital mobile telephony, and-the 4.8-kbit/s U.S. Federal Standard 1016 CELP coder, respectively, in the error-free condition.>
Satoshi Miki, Kazunori Mano, Takehiro Moriya, Kumiko Oguchi, Hitoshi Ohmuro
ICASSP (2)3
1994 Variable bit-rate speech coding based on PSI-CELP
Hitoshi Ohmuro, Kazunori Mano, Takehiro Moriya
ICSLP3
1993 An 8-bit/s speech coder based on conjugate structure CELP
Akitoshi Kataoka, Takehiro Moriya, Shinji Hayashi
ICASSP (2)2
1993 Pitch synchronous innovation CELP (PSI-CELP)
Satoshi Miki, Kazunori Mano, Hitoshi Ohmuro, Takehiro Moriya
EUROSPEECH4
1993 Training method of the excitation codebook for CELP
Takehiro Moriya, Satoshi Miki, Kazunori Mano, Hitoshi Ohmuro
EUROSPEECH1
1993 A unified approach to tree-structured and multistage vector quantization for noisy channels
abstract
The large encoding complexity and sensitivity to channel errors of vector quantization (VQ) are discussed. The performance of two low-complexity VQs-the tree-structured VQ (TSVQ) and the multistage VQ (MSVQ)-when used over noisy channels are analyzed. An algorithm is developed for the design of channel-matched TSVQ (CM-TSVQ) and channel-matched MSVQ (CM-MSVQ) under the squared-error criterion. Extensive numerical results are given for the correlation coefficient 0.9. Comparisons with the ordinary TSVQ and MSVQ designed for the noiseless channel show substantial improvements when the channel is very noisy. The CM-MSVQ, which can be regarded as a block-structured combined source-channel coding scheme, is compared with a block-structured tandem source-channel coding scheme (with the same block length as the CM-MSVQ). For the Gauss-Markov source, the CM-MSVQ outperforms the tandem scheme in all cases that the authors have considered. It is demonstrated that the CM-MSVQ is fairly robust to channel mismatch.>
Nam C. Phamdo, Nariman Farvardin, Takehiro Moriya
IEEE Trans. Inf. Theory3
1992 Two-Channel Conjugate Vector Quantizer for Noisy Channel Speech Coding
abstract
A two-channel conjugate vector quantizer is proposed in an attempt to reduce quantization distortion for noisy channels. In this quantization, two different codebooks are used. The encoder selects the channel code pair that generates the smallest distortion between the input and the averaged output vectors. These two codebooks are alternately trained by an iterative algorithm which is based on the generalized Lloyd algorithm. Coding experiments show that the proposed scheme has almost the same SNR as a conventional vector quantizer for an error-free channel. On the other hand, it has a significantly higher SNR than the conventional one for a 1% error rate. This scheme also has merits in computational complexity and storage requirements. The scheme is confirmed to be effective for a medium bit-rate speech waveform coder.>
Takehiro Moriya
IEEE J. Sel. Areas Commun.1
1990 4.8 kbit/s delayed decision CELP coder using tree coding
abstract
A 4.8-kb/s delayed decision code excited linear prediction (CELP) coder that uses tree coding is described. In conventional CELP coding, short-term and long-term prediction parameters as well as excitation parameters are sequentially determined. In the proposed delayed decision CELP coding, a tree coding method is utilized. The long-term prediction and excitation parameter candidates obtained in each subframe are listed as a tree and the optimum combined parameter sequences are selected to minimize global quantization distortion over the coding frame. The proposed coding method significantly increases the quality of the 4.8-kb/s CELP coder at the cost of an additional 5-ms coding delay.>
Kazunori Mano, Takehiro Moriya
ICASSP2
1990 Medium-delay 8 kbit/s speech coder based on conditional pitch prediction
Takehiro Moriya
ICSLP1
1990 Revised TC-WVQ speech coder for mobile communication system
Tomoyuki Ohya, Hirohito Suda, Toshio Miki, Shinji Uebayashi, Takehiro Moriya
ICSLP5
1989 An 8 kbit/s transform coder for noisy channels
abstract
The design of a speech transform coder which is robust against channel errors is presented. The bit rate is 8 kb/s including 1.3 kb/s redundancy bits for bit-selective error correction. The coder is based on transform coding with weighted vector quantization and two-channel conjugate vector quantization. Each scheme improves the robustness against channel errors without sacrificing the performance of the error-free case. Real-time operation can be achieved with 10 MIPS DSP and 8-kword memory with a 80-ms coding delay. The mean opinion score of the coded speech is comparable to that of 5-bit log PCM even at a 1% burst error rate.>
Takehiro Moriya, Hirohito Suda
ICASSP1
1988 Transform coding of speech using a weighted vector quantizer
abstract
A medium-band speech coder is proposed that uses a weighted vector quantization scheme in the transformed domain. The linear prediction residue is transformed and vector-quantized. In order to control the quantization errors in the transformed domain, adaptively weighted matching is used instead of conventional adaptive bit allocation. Therefore, the residual signal can be reconstructed by the decoder, even if the spectral envelope parameters are destroyed due to transmission errors. This coder is also capable of maintaining higher SNR (signal-to-noise ratio) performance than time-domain vector quantization coders for a wide range of computation complexities and bit rates. Coded speech is natural and unaffected by background noise. The mean opinion score for this coder at 7.2 kb/s is comparable to that of 5.5-bit log PCM coded speech sampled at 6.4 kHz.>
Takehiro Moriya, Masaaki Honda
IEEE J. Sel. Areas Commun.1
1987 Transform coding of speech with weighted vector quantization
abstract
A new transform coding of speech is proposed which employs a weighted vector quantization scheme. First, the characteristics of the weighted vector quantization are checked. Then based on these considerations, a frequency domain coder is designed for medium-band (4.8-9.6 kbps) speech coding. In this coding scheme, the linear prediction residue is transformed and vector quantized. In order to control the quantization errors in the frequency domain, adaptively weighted matching is employed instead of the conventional adaptive bit allocation. Therefore, the residual signal can be reconstructed by the decoder, even if the spectral envelope parameters are destroyed due to transmission errors. The coded speech is natural and unaffected by background noise and its mean opinion score at 7.2 kbps is comparable to that of 5.5-bit log PCM coded speech.
Takehiro Moriya, Masaaki Honda
ICASSP1
1986 Speech coder using phase equalization and vector quantization
abstract
A new speech processing and coding method is proposed which makes use of perceptual redundancy for slowly varying short-time phase characteristics. The method employs waveform conversion through a phase-equalizing filter, which is based on the time domain matched filter for the residue of Linear Predictive Coding (LPC). Phase-equalized speech is found to be almost perceptually equivalent and to be efficiently encoded by a two-stage quantization. In the first stage, vector quantization is performed for the pulse pattern in the time domain. In the second stage, vector-scalar quantization is applied to the spectral components using adaptive bit allocation. The proposed coder is proven to be superior to other coders both in terms of the SNR and the subjective quality. The averaged subjective quality at 9.6 kbps is comparable to that of a 6 bit log PCM.
Takehiro Moriya, Masaaki Honda
ICASSP1