Deepen Sinha

dblp:74/1001 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-authorArtificial intelligence and machine learning · 1Computer networks · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Speech recognition and synthesis · 100%
Theoretical computer science
1 paper
Information theory · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › speech coding
low-bit-rate speech coding
0.011994
Speech data compression through sparse coding of innovations · IEEE Trans. Speech Audio Process. 1994
Natural language and speech › Speech recognition and synthesis
speech coding
0.011994
Speech data compression through sparse coding of innovations · IEEE Trans. Speech Audio Process. 1994
Information theory › signal processing › time-frequency analysis
wavelet transform
0.011992
On the optimal choice of a wavelet for signal representation · IEEE Trans. Inf. Theory 1992

Methods — techniques the papers use, named apart from their topics

linear prediction · 0.0kalman estimation · 0.0constrained optimization · 0.0
YearPublicationVenuePosition
2020 Impact of a Shift-Invariant Harmonic Phase Model in Fully Parametric Harmonic Voice Representation and Time/Frequency Synthesis
abstract
Harmonic representation models are widely used, notably in speech coding and synthesis. In this paper, we describe two fully parametric harmonic representation and signal reconstruction alternatives that rely on a shift-invariant harmonic phase model and that implement accurate frame-based synthesis in the frequency-domain, and accurate pitch pulse-based synthesis in the time-domain. We use natural spoken and sung voice signals in order to assess the objective and subjective quality of both alternatives when parameters are exact, and when they are replaced by compact and shift-invariant harmonic phase and magnitude approximation models. We highlight the flexibility of these models and present results indicating that not only does the compact shift-invariant phase model cause a smaller impact than that caused by harmonic magnitude modeling, but it also compares favorably to results presented in the literature.
Aníbal J. S. Ferreira, Francisca Brito, Deepen Sinha
ICASSP4
2001 Multi-stream transmission for spectrally and spatially varying interference AM-DAB channels
abstract
Hybrid in band on channel (IBOC) digital audio broadcasting (DAB) simultaneously with analog amplitude modulation (AM) has been proposed as a hybrid solution to digital audio broadcasting in the AM band. Adding digital transmission in the crowded AM band is a challenging proposition. To achieve FM like audio quality, an audio coder rate of 32-64 kbit/s may be required. One of the currently proposed hybrid IBOC-AM systems is 30 kHz wide. Severe second adjacent interference may occur in certain geographical areas. For coping with such harsh transmission conditions, we present a solution based on embedded/multi-descriptive audio coding with matched multistream transmission in separate frequency bands. With loss of one frequency band, the embedded system blends to a lower audio coder rate with still a better quality than analog AM. The non-embedded system without multi-stream transmission fails catastrophically causing a severe discontinuity in quality while blending directly to analog AM.
Carl-Erik W. Sundberg, Hui-Ling Lou, Deepen Sinha
ICC3
1999 Unequal error protection methods for perceptual audio coders
abstract
In most source coded bit streams certain bits can be much more sensitive to transmission errors than others. Unequal error protection (UEP) offers a mechanism for matching error protection capability to sensitivity to transmission errors. A UEP system typically has the same average transmission rate as a corresponding equal error protection (EEP) system but offers an improved perceived signal quality at equal channel signal to noise ratio. In this work we introduce methods of UEP to the perceptual audio coder (PAC). An error sensitivity classifier divides the bits in classes of different sensitivity. Different channel codes are then applied to each class. We show how punctured convolutional codes can be used for UEP of the PAC bitstream. Experimental results for channels with uniform as well as non-uniform noise/interference level indicate that the systems with UEP exhibit graceful degradation and extended range for applications such as digital audio broadcasting (DAB).
Deepen Sinha, Carl-Erik W. Sundberg
ICASSP1
1996 Audio compression at low bit rates using a signal adaptive switched filterbank
abstract
A perceptual audio coder typically consists of a filter-bank which breaks the signal into its frequency components. These components are then quantized using a perceptual masking model. Previous efforts have indicated that a high resolution filter-bank, e.g., the modified discrete cosine transform (MDCT) with 1024 subbands, is able to minimize the bit rate requirements for most of the music samples. The high resolution MDCT, however, is not suitable for the encoding of non-stationary segments of music. A long/short resolution or "window" switching scheme has been employed to overcome this problem but it has certain inherent disadvantages which become prominent at lower bit rates (<64 kbps for stereo). We propose a novel switched filter-bank scheme which switches between a MDCT and a wavelet filter-bank based on the signal characteristics. A tree structured wavelet filter-bank with properly designed filters offers natural advantages for the representation of non-stationary segments such as attacks. Furthermore, it allows for the optimum exploitation of perceptual irrelevancies.
Deepen Sinha, James D. Johnston
ICASSP1
1994 Speech data compression through sparse coding of innovations
abstract
A new scheme for coding speech at low bit rates (4.8-16 kb/s) but still maintaining high quality is described. Speech is regarded as a piecewise-stationary random signal and its synthesis is accomplished by means of a Kalman estimator at the decoder. The Kalman estimator requires for its operation a signal model and a sequence of measurements of the states of the model. A two-stage, time-varying, all-pole filter excited by white noise is used as the speech signal model. Linear combinations of speech samples taken at sparse but periodic intervals and provided in the form of innovations serve as measurements. The role of the encoder in the proposed scheme is seen as that of extracting the signal model parameters as well as forming the measurements and transmitting this information to the decoder. An optimum measurement strategy is developed for the estimator. A procedure for shaping the error spectrum of the synthesized speech is also described. Simulation studies show that coders based on the proposed scheme can provide high-quality speech at low bit rates. Important implementation details of such coders as well as their performance results for different choices of coder parameters are given.>
Tenkasi V. Ramabadran, Deepen Sinha
IEEE Trans. Speech Audio Process.2
1993 Low bit rate transparent audio compression using a dynamic dictionary and optimized wavelets
Deepen Sinha, Ahmed H. Tewfik
ICASSP (1)1
1992 Synthesis/coding of audio signals using optimized wavelets
abstract
A novel audio synthesis/coding approach based on an optimization of the wavelet transform of the audio signal is presented. In each frame of the audio signal, the authors identify the wavelet which yields an approximation to the audio segment that requires the minimum number of bits. The wavelet approximation is constrained to have no perceptual distortion or a distortion that is less than some fixed maximum level. Two wavelet coefficient representations are considered. In the first representation, the quantization step is fixed and the magnitude of the transform coefficients is optimized. In the second approach the wavelet coefficients are adaptively quantized. The authors show that the second representation leads to simpler optimization problems. They suggest several approximations which reduce the complexity of the optimization problem. Experiments indicate that high quality speech and audio reconstruction is possible in the range of 8-9 kb/s.>
Deepen Sinha, Ahmed H. Tewfik
ICASSP1
1992 On the optimal choice of a wavelet for signal representation
abstract
Two techniques for finding the discrete orthogonal wavelet of support less than or equal to some given integer that leads to the best approximation to a given finite support signal up to a desired scale are presented. The techniques are based on optimizing certain cost functions. The first technique consists of minimizing an upper bound that is derived on the L/sub 2/ norm of error in approximating the signal up to the desired scale. It is shown that a solution to the problem of minimizing that bound does exist and it is explained how the constrained minimization over the parameters that define discrete finite support orthogonal wavelets can be turned into an unconstrained one. The second technique is based on maximizing an approximation to the norm of the projection of the signal on the space spanned by translates and dilates of the analyzing discrete orthogonal wavelet up to the desired scale. Both techniques can be implemented much faster than the optimization of the L/sub 2/ norm of either the approximation to the given signal up to the desired scale or that of the error in that approximation.>
Ahmed H. Tewfik, Deepen Sinha, Paul Jorgensen 0002
IEEE Trans. Inf. Theory2