Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Juan Carlos De Martin

dblp:89/6365 · DBLP profile ↗
← Back
58ranked-venue papers
5as first author
1since 2021 · last 2026
0000-0002-7867-1926ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 4 first-authorComputer networks · 10Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
4 papers
Content delivery and video streaming · 57% Internet architecture and protocols · 23% Wireless networking · 14%
Theoretical computer science
1 paper
Coding theory · 100%
Network and information security
1 paper
Privacy and data protection · 50% Cryptographic primitives and cryptanalysis · 50%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Content delivery and video streaming › quality of experience
video quality
0.112009
Optimized H.264 Video Encoding and Packetization for Video Transmission Over Pipeline Forwarding Networks · IEEE Trans. Multim. 2009
Content delivery and video streaming
video transmission
0.112009
Optimized H.264 Video Encoding and Packetization for Video Transmission Over Pipeline Forwarding Networks · IEEE Trans. Multim. 2009
Content delivery and video streaming
adaptive video streaming
0.112008
Variable Time Scale Multimedia Streaming Over IP Networks · IEEE Trans. Multim. 2008
Content delivery and video streaming › video transmission
multipath streaming
0.112008
Media Streaming With Network Diversity · Proc. IEEE 2008
Internet architecture and protocols
overlay networks
0.112008
Media Streaming With Network Diversity · Proc. IEEE 2008
Content delivery and video streaming › peer-assisted content distribution
peer-assisted streaming
0.112008
Media Streaming With Network Diversity · Proc. IEEE 2008
Wireless networking › link adaptation
rate adaptation
0.112008
Variable Time Scale Multimedia Streaming Over IP Networks · IEEE Trans. Multim. 2008
Internet architecture and protocols
quality of service
0.112006
Distortion-aware video communication with pipeline forwarding · ACM Multimedia 2006
Coding theory
joint source-channel coding
0.012004
Joint source-channel decoding of predictively and nonpredictively encoded sources: a two-stage estimation approach · IEEE Trans. Commun. 2004
Coding theory › error-correcting codes › decoding › decoding algorithms
joint source-channel decoding
0.012004
Joint source-channel decoding of predictively and nonpredictively encoded sources: a two-stage estimation approach · IEEE Trans. Commun. 2004
Routing and switching › switching
pipeline forwarding
0.022009
Optimized H.264 Video Encoding and Packetization for Video Transmission Over Pipeline Forwarding Networks · IEEE Trans. Multim. 2009
Distortion-aware video communication with pipeline forwarding · ACM Multimedia 2006
Audio and music processing
speech coding
0.012002
Perception-based partial encryption of compressed speech · IEEE Trans. Speech Audio Process. 2002
Cryptographic primitives and cryptanalysis › encryption
partial encryption
0.012002
Perception-based partial encryption of compressed speech · IEEE Trans. Speech Audio Process. 2002
Privacy and data protection › data confidentiality › content privacy › multimedia privacy
speech privacy
0.012002
Perception-based partial encryption of compressed speech · IEEE Trans. Speech Audio Process. 2002
Internet architecture and protocols › quality of service › delay guarantee
end-to-end delay bounds
0.012009
Optimized H.264 Video Encoding and Packetization for Video Transmission Over Pipeline Forwarding Networks · IEEE Trans. Multim. 2009
Wireless networking
wireless mesh network
0.012008
Media Streaming With Network Diversity · Proc. IEEE 2008
Coding theory › source coding › predictive coding
differential pulse-code modulation
0.012004
Joint source-channel decoding of predictively and nonpredictively encoded sources: a two-stage estimation approach · IEEE Trans. Commun. 2004
Coding theory › source coding
predictive coding
0.012004
Joint source-channel decoding of predictively and nonpredictively encoded sources: a two-stage estimation approach · IEEE Trans. Commun. 2004

Methods — techniques the papers use, named apart from their topics

packet scheduling · 0.1packetization · 0.1distortion-optimized macroblock grouping · 0.1routing · 0.1network coding · 0.1TCP-friendly rate control · 0.1PSNR analysis · 0.1perceptual bit classification · 0.1g.729 codec · 0.1simulation · 0.1pipeline forwarding · 0.1two-stage estimation · 0.0least-squares estimation · 0.0
YearPublicationVenuePosition
2026 Quantifying Privacy Risks in Synthetic Data: A Study on Black-Box Membership Inference
Giacomo Fantino, Marco Rondina, Antonio Vetrò, Juan Carlos De Martin
FASE4
2017 Removing Barriers to Transparency: A Case Study on the Use of Semantic Technologies to Tackle Procurement Data Inconsistency
Giuseppe Futia, Alessio Melandri, Antonio Vetrò, Federico Morando, Juan Carlos De Martin
ESWC (1)5
2016 Blockchain for the Internet of Things: A systematic literature review
abstract
In the Internet of Things (IoT) scenario, the block-chain and, in general, Peer-to-Peer approaches could play an important role in the development of decentralized and dataintensive applications running on billion of devices, preserving the privacy of the users. Our research goal is to understand whether the blockchain and Peer-to-Peer approaches can be employed to foster a decentralized and private-by-design IoT. As a first step in our research process, we conducted a Systematic Literature Review on the blockchain to gather knowledge on the current uses of this technology and to document its current degree of integrity, anonymity and adaptability. We found 18 use cases of blockchain in the literature. Four of these use cases are explicitly designed for IoT. We also found some use cases that are designed for a private-by-design data management. We also found several issues in the integrity, anonymity and adaptability. Regarding anonymity, we found that in the blockchain only pseudonymity is guaranteed. Regarding adaptability and integrity, we discovered that the integrity of the blockchain largely depends on the high difficulty of the Proof-of-Work and on the large number of honest miners, but at the same time a difficult Proof-of-Work limits the adaptability. We documented and categorized the current uses of the blockchain, and provided a few recommendations for future work to address the above-mentioned issues.
Marco Conoscenti, Antonio Vetrò, Juan Carlos De Martin
AICCSA3
2015 On the effects of sender-receiver concealment mismatch on multimedia communication optimization
Enrico Masala, Fabio De Vito, Juan Carlos De Martin
Multim. Tools Appl.3
2014 Measuring DASH streaming performance from the end users perspective using neubot
abstract
The popularity of DASH streaming is rapidly increasing and a number of commercial streaming services are adopting this new standard. While the benefits of building streaming services on top of the HTTP protocol are clear, further work is still necessary to evaluate and enhance the system performance from the perspective of the end user. Here we present a novel framework to evaluate the performance of rate-adaptation algorithms for DASH streaming using network measurements collected from more than a thousand Internet clients. Data, which have been made publicly available, are collected by a DASH module built on top of Neubot, an open source tool for the collection of network measurements. Some examples about the possible usage of the collected data are given, ranging from simple analysis and performance comparisons of download speeds to the performance simulation of alternative adaptation strategies using, e.g., the instantaneous available bandwidth values.
Simone Basso, Antonio Servetti, Enrico Masala, Juan Carlos De Martin
MMSys4
2013 Challenges and Issues on Collecting and Analyzing Large Volumes of Network Data Measurements
Enrico Masala, Antonio Servetti, Simone Basso, Juan Carlos De Martin
ADBIS (2)4
2012 MDL-based joint denoising and compression of intracortical signals
abstract
Intra-cortical signals are usually affected by high levels of noise (0 dB SNR is not uncommon) either due to the recording equipment or to magnetical and electrical couplings between surrounding sources and the recording system. Besides from hindering effective exploitation of the information content in the signals, noise also influences the bandwidth needed to transmit them, which is a problem especially when a large number of channels are to be recorded. In this paper we propose a novel technique for joint denoising and compression of intra-cortical signals based on the Minimum Description Length principle (MDL). This method was tested on simulated signals and the results showed that the proposed technique achieves improvements in SNR (up to .6 dB over MNML for very noisy signals) and compression ratios greater than alternative denoising/compression methods.
Elias S. G. Carotti, Winnie Jensen, Juan Carlos De Martin, Dario Farina
ICASSP3
2011 The network neutrality bot architecture: A preliminary approach for self-monitoring of Internet access QoS
abstract
The “network neutrality bot” (Neubot) is an evolving software architecture for distributed Internet access quality and network neutrality measurements. The core of this architecture is an open-source agent that ordinary users may install on their computers to gain a deeper understanding of their Internet connections. The agent periodically monitors the quality of service provided to the user, running background active transmission tests that emulate different application-level protocols. The results are then collected on a central server and made publicly available to allow constant monitoring of the state of the Internet by interested parties. In this article we describe how we enhanced Neubot architecture both to deploy a distributed broadband speed test and to allow the development of plug-in transmission tests. In addition, we start a preliminary discussion on the results we have collected in the first three months after the first public release of the software.
Simone Basso, Antonio Servetti, Juan Carlos De Martin
ISCC3
2010 Content-adaptive traffic prioritization of spatio-temporal scalable video for robust communications over QoS-provisioned 802.11e networks
Attilio Fiandrotti, Dario Gallucci, Enrico Masala, Juan Carlos De Martin
Signal Process. Image Commun.4
2009 Content-Adaptive Robust H.264/SVC Video Communications over 802.11e Networks
abstract
In this paper we present a low-complexity traffic prioritization strategy for video transmission using the H.264 scalable video coding (SVC) standard over 802.11e wireless networks.The first part of this work focuses on assessing the perceptual impact of data loss in the various enhancement layers using a wide set of H.264/SVC encoded videos.The analysis shows that perceptual impairments are highly correlated with the motion activity in the video sequence.Thus, we propose an adaptive unequal error protection strategy which identifies the most perceptually important parts of the enhancement layers in the video sequence by means of a low complexity macroblock motion analysis process.The algorithm is tested by simulating a realistic 802.11e based home network scenario.Results obtained on a large set of video sequences show that the proposed content-aware traffic prioritization strategy enables PSNR gains up to 2.5 dB as well as noticeable visual quality improvements with respect to a traditional prioritization strategy aiming at minimizing error propagation.
Dario Gallucci, Attilio Fiandrotti, Enrico Masala, Juan Carlos De Martin
AINA4
2009 Low-complexity perceptual packet marking for speech transmission over tiny mote device
abstract
Multimedia applications in Wireless Sensor Networks (WSNs) require different approaches with respect to traditional networks due to the adopted devices' limitations in energy consumption and computational capabilities. This paper proposes a low-complexity algorithm for selecting speech packets according to their perceptual importance for voice transmission over WSNs. The proposed algorithm allows wireless sensor nodes to select which packets to protect during transmission in order to increase speech quality while at the same time minimizing the necessary energy. Experimental results based on cooperative transmission among devices show that the proposed algorithm achieves a good speech quality level while reducing the need to protect packets by 40% when compared to random selection.
Matteo Petracca, Juan Carlos De Martin, Gustavo Litovsky, Marco Tacca, Andrea Fumagalli
ICME2
2009 Text-independent compressed domain speaker verification for digital communication networks call monitoring
abstract
In this paper we present a text-independent automatic speaker verification system that works in the compressed domain using GSM AMR coded speech. While traditional approaches process LPC-based cepstral coefficients extracted from LPC related bitstream coefficients, our objective is to study the feasibility of a system that directly processes the raw bitstream transmitted over a digital communication network. In a text-independent closet-set task, with a database of 100 speakers, the proposed system achieves an EER equal to 5.93%, 4.41% and 3.80% for 10, 20 and 30 second long test speech segments respectively.
Matteo Petracca, Antonio Servetti, Juan Carlos De Martin
ICME3
2009 Perceptual based voice multi-hop transmission over wireless sensor networks
abstract
Multimedia applications over wireless sensor networks (WSNs) are rapidly gaining interest by the research community in order to develop new and mission critical services such as environmental video monitoring and emergency speech calls. In this work we analyze the possibility of sending voice using a network of wireless tiny motes with the final goal of enhancing speech quality by protecting the most perceptually important packets. We first evaluate the speech quality for a modified version of the ITU-T G.711 standard implemented to fit the particular selected hardware. Hence, we propose a low-complexity measure to evaluate the perceptual importance of speech packets. When performing single-hop experimental data collection, we apply packet redundancy (protection) by using a cooperative mote (relay) which retransmits speech packets that are perceptually important to protect them against potential transmission losses. Collected experimental results are then used to assess multi-hop performance, showing that the combination of the selected hardware and the proposed perceptual marking algorithm achieves good speech quality levels, according to the MOS scale, while reducing the percentage of protected packets by 40% when compared to random protection.
Matteo Petracca, Gustavo Litovsky, Alessandro Rinotti, Marco Tacca, Juan Carlos De Martin, Andrea Fumagalli
ISCC5
2009 Optimized H.264 Video Encoding and Packetization for Video Transmission Over Pipeline Forwarding Networks
abstract
Previous works showed that the quality-of-service (QoS) requirements of multimedia applications can be optimally satisfied by pipeline forwarding (PF) by providing end-to-end delay guarantees as well as high network resource utilization. However, the unavoidable mismatch between reserved resources and the unpredictable traffic profile of a video stream has an impact on the resulting application layer quality. Therefore, a new low-complexity H.264 video encoding and packetization scheme based on a distortion-optimized macroblock grouping technique is designed here to maximize the performance of video transmission on PF networks. The scheme considers the perceptual importance of the different parts of the video data to group the most important information in few packets that are the natural candidates to receive the deterministic service provided by PF. Results show peak signal-to-noise ratio (PSNR) gains up to 2.5 dB over traditional video encoding and packetization schemes, as well as more graceful degradation in case of high network load.
Enrico Masala, Andrea Vesco, Mario Baldi, Juan Carlos De Martin
IEEE Trans. Multim.4
2008 Matrix-based linear predictive compression of multi-channel surface emg signals
abstract
We propose a linear predictive coding technique for multichannel electromyographic (EMG) recordings. The signals are acquired using two-dimensional grid of electrodes which generate strongly correlated signals. Previous work only considered spectral redundancy across the signal matrix. In this paper we exploit the correlation present in the residual signals, i.e., the signals after the short term prediction. The proposed technique achieves a compression ratio of about 1divide9, i.e., slightly better than spectral-only decorrelation methods, but with a strong increase of approximately 3.2 dB SNR in the quality of the reconstructed waveform.
Elias S. G. Carotti, Juan Carlos De Martin, Roberto Merletti, Dario Farina
ICASSP2
2008 The Neubot project: A collaborative approach to measuring internet neutrality
abstract
The Internet was designed to be neutral with respect to kinds of applications, senders and destinations. Such design choice made very fast packet switching possible, while preserving, at the same time, strong openness towards unforeseen uses of the Internet Protocol. The result has been an extraordinary outburst of innovation, as well as a level playing field for citizens, associations and companies worldwide. With the advent of ldquodeep packet inspectionrdquo technology, however, fine-grained discrimination of Internet flows is now possible, be that for economic or other reasons. Collecting quantitative data on the behavior of telecommunications providers with respect to traffic discrimination thus becomes crucial, particularly at a time when policy changes are widely discussed. The ldquoNetwork Neutrality Botrdquo (Neubot) project is based on a lightweight, open source computer program, the Neubot, that, downloaded and installed by Internet users, performs distributed measurements of the traffic characteristics of segments of the global Internet. The collected data will allow constant monitoring of the actual state of the Internet, enabling both a deeper understanding of such crucial infrastructure and a more reliable basis for discussing network neutrality policies.
Juan Carlos De Martin, Andrea Glorioso
ISTAS1
2008 Media Streaming With Network Diversity
abstract
Today's packet networks including the Internet offer an intrinsic diversity for media distribution in terms of available network paths and servers or information sources. Novel communication infrastructures such as ad hoc or wireless mesh networks use network diversity to extend their reach at low cost. Diversity can bring interesting benefits in supporting resource greedy applications such as media streaming services, by aggregation of bandwidth and computing resources. Typically, overlay network architectures compensate for lack of quality-of-service guarantees in the network by introducing redundancy in the media delivery system through network diversity. They can support efficient multimedia services when routing, coding, and scheduling algorithms are able to adapt to both the media information and the dynamic network status. This paper presents an overview of the distributed streaming solutions that profit from network diversity in order to improve the quality of multimedia applications. We discuss the coding techniques used for adaptive and flexible media streaming with network diversity. We describe the problem of media streaming with path diversity and focus on routing, path computation, and packet scheduling problems in multipath networks. Then, the advantages of server or source peer diversity in collaborative streaming solutions are discussed. Lastly, we present an overview of wireless mesh networks and focus on the typical constraints imposed by these novel communication models on media streaming with network diversity.
Pascal Frossard, Juan Carlos De Martin, M. Reha Civanlar
Proc. IEEE2
2008 Variable Time Scale Multimedia Streaming Over IP Networks
abstract
This paper presents a comprehensive analysis of a variable time-scale streaming technique, VTSS, according to which rate changes are obtained by varying the inter-packet transmission interval, rather than altering, as in most cases, the source coding rate. Instead of constraining the transmitter to operate in real-time, the time scale of the packet scheduler can vary between zero, when the network is congested, to as faster than real-time as the channel bandwidth allows, when the network is lightly loaded. Although this approach is reportedly used in commercial streaming products, so far the technique has not yet been analyzed in a rigorous fashion, nor it has been compared to other state-of-the-art streaming techniques. This work first presents a theoretical analysis of the performance achievable by the VTSS approach, and it shows that, for the same channel conditions, VTSS yields a total distortion which is lower or, in the worst case, equal than the distortion of the standard real-time source-rate adaptive approach. A lower bound on receiver buffer size is also derived. Network simulations then analyze the performance of a TCP-friendly test implementation of VTSS compared with an ideal real-time source rate-adaptive technique, whose performance, being ideal, represents the upper bound of any transmission scheme based on source rate adaptation. The simulation results, also based on actual network traces, show that the VTSS approach delivers higher perceptual quality (up to 1.2 dB PSNR in the considered scenarios) and reduced video quality fluctuations (1.6 dB standard deviation PSNR, instead of 4.9 dB) for a wide range of standard video sequences. Perceptual quality evaluation by means of PVQM confirms such results. The gains, as expected, are even more pronounced (7.6 dB PSNR on average) if compared to real-time constant bit-rate video transmission.
Enrico Masala, Davide Quaglia, Juan Carlos De Martin
IEEE Trans. Multim.3
2007 ACELP-Based Compression of Multi-Channel Surface EMG Signals
abstract
In this paper we extend a lossy compression technique for surface EMG signals, which is based on the algebraic code excited linear prediction (ACELP) paradigm, to compress multi-channel surface EMG recordings by exploiting the correlation between the line spectral frequencies (LSF). Experimental results show that the LSFs of the inner signals in a multi-channel recording can be efficiently represented with 13 bit/frame, versus the 38 bit/frame needed by independent ACELP coding of each signal, thus saving 66% of the bandwidth needed to transmit these coefficients while maintaining comparable performance in terms of the SNR, average rectified value and root mean square of the waveform, and mean and median frequencies of the power spectrum.
Elias S. G. Carotti, Juan Carlos De Martin, Roberto Merletti, Dario Farina
ICASSP (2)2
2007 Performance Analysis of Distributed Speech Recognition Services over Noisy 802.11b Wireless Networks
abstract
The performance of an AURORA-like distributed speech recognition system over IEEE 802.11 WLANs is studied. The recognition features are packetized and sent over an 802.11b network. At the receiver recognition is performed. Two different scenarios are simulated to analyze DSR performance in presence of losses due to either low received power or to network congestion. Varying recognizer complexities, packet lengths, number of concurrent flows, and signal power levels are considered in both scenarios. Experimental results on a connected digits task show that for low signal power levels, the best recognition performance is obtained when speech features are sent in small IP packets, while in the case of network congestion the best performance is obtained by increasing the packet size.
Alessandro Rinotti, Piero Demichelis, Juan Carlos De Martin
ISM3
2006 Compression of Surface Emg Signalswith Algebraic Code Excited Linear Prediction
abstract
In this paper we investigate a lossy coding technique for surface EMG signals which is based on the algebraic code excited linear prediction (ACELP) paradigm, widely used for speech signal coding. The algorithm was adapted to the EMG characteristics and tested on both simulated and experimental signals. A fixed compression ratio of 87.3% was chosen. On simulated signals, the mean square error in signal reconstruction and the percentage error in average rectified value after compression were 10.43 % and 5.52 %, respectively. On experimental signals, they were 6.74% and 3.11%. The mean power spectral frequency and third order power spectral moment were estimated with relative error smaller than 1.36% and 1.70%, respectively, for simulated signals, and 3.74% and 2.28% for experimental signals. It was concluded that the proposed coding scheme can be effectively used for high rate, low distortion and low-delay compression of surface EMG signals
Elias S. G. Carotti, Juan Carlos De Martin, Roberto Merletti, Dario Farina
ICASSP (3)2
2006 An Analysis of Constant Bitrate and Constant PSNR Video Encoding for Wireless Networks
abstract
In wireless networks, transmission of constant-quality, high bitrate video is a challenging task due to channel capacity and buffer limitations. Content adaptive rate control, is used as a solution to this problem. Instead of transmitting all of the video content at low quality, the most important content can be transmitted at high quality while still preserving an acceptable quality for the remaining segments. Furthermore, the rate control strategy inside the individual temporal segments plays a key role for the network performance and viewing quality. Although constant quality video encoding inside the temporal segments is preferable for the best viewing experience, it causes more network packet losses due to adverse bitrate fluctuations in the video stream. In cases when the network is too much loaded, it may be better to employ constant bitrate encoding for network friendliness. In this paper, a performance analysis of constant bitrate and constant peak signal-to-noise ratio encoding for content adaptive rate controlled video streaming over wireless networks is presented. Experimental results obtained using AVC/H.264 encoding in a CDMA/HDR multi-user environment with cross-layer optimized scheduling show performance comparisons of CBR and CPSNR encoding.
Tanir Ozcelebi, Fabio De Vito, A. Murat Tekalp, M. Reha Civanlar, M. Oguz Sunay, Juan Carlos De Martin
ICC6
2006 Performance Analysis of Compressed-Domain Automatic Speaker Recognition as a Function of Speech Coding Technique and Bit Rate
abstract
Compressed-domain automatic speaker recognition is based on the analysis of the compressed parameters of speech coders. The objective is to perform low-complexity on-line speaker recognition for VoIP in the compressed domain, without the need to decode or resynthesize the speech bitstream. In this paper, we present initial results in determining the recognition accuracy that can be achieved with five widely used speech coding standards. Experiments with a database of 14 speakers obtain a recognition ratio close to 100% after the analysis of 30 seconds of active speech for most of the considered speech coders and rates. In particular, the results show that performance does not strictly depend on coding rate or codec speech quality
Matteo Petracca, Antonio Servetti, Juan Carlos De Martin
ICME3
2006 Distortion-aware video communication with pipeline forwarding
abstract
This paper tackles the issue of optimizing the transport of video over packet networks with respect to both resource utilization and user perceived quality. Previous work showed that the quality of service requirements of multimedia applications can be satisfied by pipeline forwarding of packets. However, the current Internet is not based on such technology and its incremental introduction raises questions on how to handle video packets generated by pipeline forwarding unaware sources at the interface between a subnetwork deploying conventional packet scheduling techniques and one implementing pipeline forwarding. This work proposes to use the perceptual importance of the carried video samples to determine which packets shall be transferred with pipeline forwarding - thus receiving deterministic service - and which with a traditional, e.g., best effort or differentiated, service. Simulation results with the first implemented variants of this solution are presented.
Mario Baldi, Juan Carlos De Martin, Enrico Masala, Andrea Vesco
ACM Multimedia2
2005 Linear predictive coding of myoelectric signals
abstract
Despite the great interest towards long term recordings of electromyographic (EMG) signals, which find applications, for example, in telemedicine, only a few studies have dealt with the compression of these signals. We propose a lossy coding technique for surface EMG signals. The technique is based on the linear predictive coding paradigm widely used for speech compression. The algorithm was tested on both simulated and experimental signals. Mean frequency, median frequency, variance, skewness and kurtosis of the EMG signals were preserved with an error less than 3% with respect to the original values for synthetic signals and experimental signals, reducing the bitrate from 24 kbit/s (12 kbit/s after downsampling) to 352 bit/s, with a compression factor of 97.1%. It was concluded that the linear predictive coding paradigm can be effectively used for high rate compression of surface EMG signals when preservation of only the power spectrum of the signal is of interest. This has applications in ergonomics and occupational medicine.
Elias S. G. Carotti, Juan Carlos De Martin, Dario Farina, Roberto Merletti
ICASSP (5)2
2005 Standard Compatible Error Correction for Multimedia Transmissions Over 802.11 WLAN
abstract
In this paper, we analyze a standard compatible error correction technique for multimedia transmissions over 802.11 WLANs that exploits, when available, the information of previous erroneous transmissions. The basic idea is to store erroneous frames for error correction purposes. More specifically, at the receiver each bit is estimated with a majority criterion. The performance of different standard compliant error recovery techniques have been evaluated using actual transmission experiments in various channel conditions. Optimal tradeoffs between complexity, memory and perceived quality have been determined studying the quality gains that can be achieved for the specific case of multimedia applications. Perceived quality has been evaluated using objective measures, e.g. ITU-T PESQ for voice and PSNR for video. Results show that the majority combining approach is particularly effective for multimedia communications, even in very noisy scenarios. Gains up to about one unit on the MOS scale for speech and up to 5-6 dB PSNR in case of video have been measured with respect to the standard ARQ technique
Enrico Masala, Antonio Servetti, Juan Carlos De Martin
ICME3
2005 Low-complexity automatic speaker recognition in the compressed GSM AMR domain
abstract
This paper presents an experimental implementation of a low-complexity speaker recognition algorithm working in the compressed speech domain. The goal is to perform speaker modeling and identification without decoding the speech bitstream to extract speaker dependent features, thus saving important system resources, for instance, in mobile devices. The compressed bitstream values of the widely used GSM AMR speech coding standard are studied to identify statistics enabling fair recognition after a few seconds of speech. Using Euclidean distance measures on elementary statistical values such as coefficient of variation and skewness of nine standard GSM AMR parameters delivers recognition accuracies close to 100% after about 20 seconds of active speech for a database of 14 speakers recorded in a normal room environment.
Matteo Petracca, Antonio Servetti, Juan Carlos De Martin
ICME3
2005 Hybrid Bitrate/PSNR Control for H.264 Video Streaming to Roaming Users
abstract
In wireless communications, the available throughput depends on several parameters, like physical layer, base station distance, fading and interference. Users experience changes in bandwidth within a cell and among same-technology cells, but also among different networks. Moreover, in case of video transmission, the user may specify a desired quality level. The source should encode the stream at a quality as close as possible to this value, without exceeding the available bitrate. We propose a technique to decide whether to encode at constant quality, if resources are enough, or at constant bitrate, if the throughput is not sufficient. With negligible complexity, it proved to obtain better PSNR/bitrate ratios, with respect to only constant bitrate and only constant PSNR coding. We also show this algorithm working in a realistic scenario of a user roaming among heterogeneous networks (WLAN and UMTS). Also in this case, the algorithm proved to achieve high quality/bitrate ratios.
Fabio De Vito, Federico Ridolfo, Juan Carlos De Martin
ISM3
2005 Performance Evaluation of H.264 Video Streaming over Inter-Vehicular 802.11 Ad Hoc Networks
abstract
This paper evaluates the performance of video streaming in inter-vehicular environments using the 802.11 ad hoc network protocol. We performed transmission experiments while driving two cars equipped with 802.11b standard devices in urban and highway scenarios. Different sequences, bitrates and packctization policies have been tested. The experiments show that each scenario presents peculiar characteristics in terms of average link availability and SNR, which can be exploited to develop more efficient applications. In this paper we also determine the best packetization policies for the two scenarios, showing that large packets lead to better performance in the highway scenario and vice versa. Perceptual quality results indicate that the best packetization policy achieves consistent gains in terms of PSNR values (up to 5 dB), and reduced quality variations, with respect to a fixed-policy transmission technique
Paolo Bucciol, Enrico Masala, Nobuo Kawaguchi, Kazuya Takeda, Juan Carlos De Martin
PIMRC5
2004 Rate-Distortion Optimized Slicing, Packetization and Coding for Error Resilient Video Transmission
abstract
This paper presents an algorithm to optimize the tradeoff between rate and expected end-to-end distortion of a video sequence transmitted over a packet network. The approach optimizes the source coding parameters, slicing, network QoS class selection and/or error control coding parameters, and accounts for the effects of compression, packetization, error propagation, and concealment at the decoder. It builds on, and substantially extends the applicability of, the recursive optimal per-pixel estimate (ROPE) technique for end-to-end distortion estimation. A trellis-based algorithm is introduced in order to overcome macroblock interdependencies in the estimation procedure, and allow adaptive slicing. Moreover, we propose a complementary packetization scheme to efficiently arrange the slices into packets for FEC protection while minimizing rate loss due to padding. Simulations demonstrate consistent gains over currently used techniques.
Enrico Masala, Kenneth Rose, Juan Carlos De Martin
Data Compression Conference4
2004 Cross-layer perceptual ARQ for H.264 video streaming over 802.11 wireless networks
abstract
We present a new cross-layer ARQ algorithm for video streaming over 802.11 wireless networks. The algorithm combines application-level information about the perceptual and temporal importance of each packet into a single priority value, which drives packet selection at each retransmission opportunity. Hence, only the most perceptually important packets are retransmitted, delivering higher perceptual quality and less bandwidth usage compared to the standard 802.11 MAC-layer ARQ scheme. H.264 video streaming based on the proposed technique has been simulated using ns in a realistic home network scenario, using the standard ARQ technique for all interfering traffic. Results show that the proposed method consistently outperforms the standard MAC-layer 802.11 retransmission scheme, delivering more than 1.5 dB PSNR gains using approximately half of the retransmission bandwidth.
Paolo Bucciol, Gabriele Davini, Enrico Masala, Enrica Filippi, Juan Carlos De Martin
GLOBECOM5
2004 Fast implementation of the MPEG-4 AAC main and low complexity decoder
abstract
We present a fast software implementation of the MPEG-4 AAC (advanced audio coding) main and low complexity (LC) decoder. The reference implementation is analyzed and selected algorithms are presented to improve the performance of most of its building blocks, i.e., bitstream de-formatter, noiseless decoding, prediction, and filterbank. The code is further optimized by means of assembler procedures for the filterbank and prediction tools to exploit the Intel Pentium streaming SIMD extension (SSE) instruction set. SSE performs operations over four single precision floating-point numbers in a single instruction using 128-bit long registers. Results show that the presented decoder proves to be almost five times faster than the reference implementation, while preserving its full compatibility with the MPEG-4 standard.
Antonio Servetti, Alessandro Rinotti, Juan Carlos De Martin
ICASSP (5)3
2004 Perceptual ARQ for H.264 video streaming over 3G wireless networks
abstract
This paper presents a new ARQ algorithm for video streaming over wireless channels. The algorithm takes into account the perceptual and temporal importance of each packet to determine the packet scheduling which maximizes the perceived quality. A simple and flexible function to combine the perceptual importance with the real-lime constraints of each packet to determine which is the best packet to transmit at each transmission opportunity is proposed. The perceptual importance is evaluated using the analysis-by-synthesis technique. The performance of the proposed algorithm has been analyzed by simulating the transmission of H.264 encoded sequences over a 144 kbit/s UMTS channel and compares the proposed method with time-driven ARQ techniques, using PSNR as distortion measure. The results show that for the considered channel conditions the proposed method delivers gains up to 2 dB with respect to the time-driven ARQ technique.
Paolo Bucciol, Enrico Masala, Juan Carlos De Martin
ICC3
2004 Model-based distortion estimation for perceptual classification of video packets
abstract
In video communications over IP networks, quality of service (QoS) guarantees must be introduced to limit the effect of packet losses. In particular, end-to-end QoS can be improved if packets are protected according to the distortion that would be introduced at the receiver by their loss. In the traditional analysis-by-synthesis (AbS) approach, each packet is assumed lost, error concealment applied, the sequence decoded, and the resulting overall distortion computed. This process produces reliable distortion estimates, but is computationally demanding. In this work, we present a hybrid approach: the distortion introduced in the current frame is evaluated with the AbS method, while the distortion in future frames is estimated by means of a statistical error-propagation model. Results obtained on eight, widely different H.264 sequences show that the proposed model successfully estimates overall distortion with very low complexity. Network simulations also show that model-based packet classification, when used for video transmission over DiffServ networks, delivers PSNR results which are consistently within 0.1 dB compared to the AbS technique.
Fabio De Vito, Davide Quaglia, Juan Carlos De Martin
MMSP3
2004 Joint source-channel decoding of predictively and nonpredictively encoded sources: a two-stage estimation approach
abstract
A common joint source-channel (JSC) decoder structure for predictively encoded sources involves first forming a JSC decoding estimate of the prediction residual and then feeding this estimate to a standard predictive decoding (synthesis) filter. In this paper, we demonstrate that in a JSC decoding context, use of this standard filter is suboptimal. In place of the standard filter, we choose the synthesis filter coefficients to give a least-squares (LS) estimate of the original source, based on given training data. For first-order differential pulse-code modulation, this yields as much as 0.65-dB gain in reconstructing first-order Gauss-Markov sources. More gains are achieved with modest additional complexity by increasing the filter order. While performance can also be enhanced by increasing the source's Markov model order and/or the decoder's lookup table memory, complexity grows exponentially in these parameters. For both predictive and nonpredictive coding, our LS approach offers a strategy for increasing the estimation accuracy of JSC decoders while retaining manageable complexity.
David J. Miller 0001, Elias S. G. Carotti, Yu-Wei Wang, Juan Carlos De Martin
IEEE Trans. Commun.4
2003 Frequency-selective partial encryption of compressed audio
abstract
The widespread adoption of compressed digital audio increasingly demands effective ways to conjugate ease of distribution with content protection. In this paper a low-complexity partial-encryption scheme for MPEG audio is presented. The aim is to provide listeners with sample-quality audio material that can be upgraded to full-quality by simply acquiring a key and decrypting a few selected bits (1-10% of the total bitstream) of the already available data. Sample-quality is obtained by limiting the frequency content of the signal, and it is achieved without altering the compatibility of the compressed audio bitstream. Audio material partially encrypted with the proposed scheme can thus be freely distributed for evaluation and then easily unlocked to achieve full quality without any further transmission of audio data.
Antonio Servetti, Cristiano Testa, Juan Carlos De Martin
ICASSP (5)3
2003 A simulative study of analysis-by-synthesis perceptual video classification and transmission over DiffServ IP networks
abstract
This paper presents the results of transmission of video data on 2-class DiffServ IP networks using perceptual packet classification and slicing. An analysis-by-synthesis technique to identify perceptually important video regions, to create optimal video slices and to assign the resulting packets to the appropriate DiffServ classes is described. The proposed technique was implemented using the ISO/IEC MPEG-2 video coding standard. Several transmission scenarios, including homogeneous video traffic and interfering FTP traffic, were simulated using network simulator (NS). The proposed perception-based video transmission approach outperformed classical data partitioning in all tested network usage and potential to match time-varying channels. Substantially higher PSNR values than the regular best-effort case were also obtained assigning to the high-QoS class as little as 10% of the traffic.
Fabio D'Agostino, Enrico Masala, Laura Farinetti, Juan Carlos De Martin
ICC4
2003 Perceptually-evaluated loss-delay controlled adaptive transmission of MPEG video over IP
abstract
This paper presents the adaptive video over IP (AViP) approach to transmit video sequences. A rate-selection algorithm based on both delay and loss indication is presented and its performance measured using actual MPEG-2 video sequences, network simulations and objective measures of perceptual quality. The results show that the AViP approach leads to efficient use of available network resources, reactiveness to congestions and TCP-friendliness. Moreover, AViP delivers significantly higher perceptual levels of quality than traditional constant-bit-rate systems operating at the same average level.
Gabriele Davini, Davide Quaglia, Juan Carlos De Martin, Claudio Casetti
ICC3
2003 Low-complexity lossless video coding via adaptive spatio-temporal prediction
abstract
Lossless coding of video sequences is becoming increasingly attractive for applications as diverse as digital cinema, medical imaging and professional video processing. We present a new low-delay, low-complexity algorithm for lossless color video compression. The proposed technique adaptively exploits temporal, spatial and spectral redundancy. Key features of the coder are a backward-adaptive temporal predictor, an intra-frame spatial predictor and adaptive optimal weighting of both predictive components. The residual error is entropy coded by a context-based arithmetic encoder. The proposed lossless video encoder delivers significantly higher compression gains than traditional approaches at approximately the same complexity and delay, enabling efficient storage and communications for emerging lossless video applications.
Elias S. G. Carotti, Juan Carlos De Martin, Angelo Raffaele Meo
ICIP (2)2
2003 Analysis-by-synthesis distortion computation for rate-distortion optimized multimedia streaming
abstract
This paper presents an analysis-by-synthesis technique to evaluate the perceptual importance of multimedia packets for rate-distortion optimized streaming. The proposed technique, instead of relying on a priori information, computes the distortion that would be caused by the loss of each single packet, including the effects of error propagation and receiver-side error concealment. A rate-distortion optimized streaming algorithm is presented to compare the perceptual performance obtained using content-adaptive analysis-by-synthesis distortion values versus distortion values obtained using a priori knowledge of the statistical importance of the elements of the compressed multimedia bitstream. Simulations with video test sequences compressed with the MPEG-2 coding standard show that the proposed technique delivers substantial and consistent PSNR gains (1.2-2.8 dB) with respect to ideal frame type-driven a priori distortion evaluation for a wide range of channel conditions. Compared to distortion-agnostic streaming techniques such as SoftARQ, the gain is even more pronounced.
Enrico Masala, Juan Carlos De Martin
ICME2
2003 Adaptive packet classification for constant perceptual quality of service delivery of video streams over time-varying networks
abstract
This paper describes a technique to deliver video streams with constant perceptual quality of service (QoS) over time-varying packet erasure channels that support differentiated classes of service. During compression, the encoder estimates the distortion introduced at the decoder by each video packet both in case it is received and in case it is lost. Concealment and error propagation due to inter-frame prediction are taken into account. During transmission, an optimization algorithm assigns packets to different service classes according to the estimated distortion, the current channel status and a constraint on the desired quality at the receiver. This technique is compared with other video packet classification approaches in the specific case of a DiffServ IP network implementing the assured forwarding scheme. Network simulations show that the proposed technique delivers higher and more constant levels of perceptual QoS than traditional approaches. Moreover, the technique is characterized by reactiveness to congestion and fairness in the use of network resources.
Davide Quaglia, Juan Carlos De Martin
ICME2
2002 Backward-adaptive lossless compression of video sequences
abstract
We present our new low-complexity compression algorithm for lossless coding of video sequences. This new coder produces better compression ratios than lossless compression of individual images by exploiting temporal as well as spatial and spectral redundancy. Key features of the coder are a pixel-neighborhood backward-adaptive temporal predictor, an intra-frame spatial predictor and a differential coding scheme of the spectral components. The residual error is entropy coded by a context-based arithmetic encoder. This new lossless video encoder outperforms state-of-the-art lossless image compression techniques, enabling more efficient video storage and communications.
Elias S. G. Carotti, Juan Carlos De Martin, Angelo Raffaele Meo
ICASSP2
2002 Interactive DSP educational platform for real-time subband audio coding
abstract
We present an interactive educational DSP platform for realtime subband audio coding. The tool integrates a DSP course on advanced audio coding and focuses on wavelets, subband coding, and signal quantization. The package consists of a Windows control application running on a PC and real-time DSP code implemented on a Texas Instruments DSP Starter Kit (DSK), which performs the actual audio encoding and decoding process. The tool is interactive since students can customize the codec by changing some encoding parameters through the PC interface and then verify the effect on the coding process. The DSP codec accepts audio samples in real-time via the DSK audio analog input port and sends the reconstructed samples to the output port. The graphical user interface allows to choose the type of filter-bank and scalar quantization; the availability of the full source code and of a DSK makes possible to experience first-hand the main aspects of real-world DSP development, e.g., real-time programming, arithmetic precision and algorithmic delay. The software is available at http://multimedia.polito.it/dspwavelab.
Davide Quaglia, Alfonso Montuori, Eros Pasero, Juan Carlos De Martin
ICASSP4
2002 Performance analysis of Distributed Speech Recognition over IP networks on the AURORA database
abstract
We present results on the performance of Distributed Speech Recognition operating over simulated IP networks. ETSI AURORA front-end running at client nodes extracts the speech parameters, packetizes and sends them as real-time IP traffic to a remote recognizer based on Continuous Density Hidden Markov Models. The experimental framework is the ETSI STQ-AURORA Project Database 2.0. The impact of transmission over IP networks is modeled by (1) random losses, (2) losses generated by a Gilbert model and (3) network simulations. Results show that random losses and moderately bursty losses do not significantly affect the recognition performance. Strongly bursty packet losses, as those generated by real-time and Web traffic competing over a network bottleneck, instead, can have a very negative impact on recognition performance, indicating that DSR over the Internet, to be successful, requires high levels of Quality of Service.
Daniele Quercia, Laura Docío Fernández, Carmen García-Mateo, Laura Farinetti, Juan Carlos De Martin
ICASSP5
2002 Perception-based selective encryption of G.729 speech
abstract
Mobile multimedia applications, the focus of many forth-coming wireless services, increasingly demand low-power techniques implementing content protection and customer privacy. In this paper a low-complexity, perception-based partial encryption scheme for telephone-bandwidth speech is presented. Speech compressed by a widely-used speech coding algorithm, ITU-T G.729 CS-ACELP at 8 kb/s, is partitioned in two classes, one, the most perceptually relevant, to be encrypted, the other, to be left unprotected. Encryption of about 45% of the bitstream achieves content protection equivalent to full encryption of the bitstream, as verified by both objective measures and formal listening tests. Low-power, portable devices can, therefore, implement very high levels of speech-content protection at a fraction of the computational load of current techniques, freeing resources for other tasks and enabling longer battery life.
Antonio Servetti, Juan Carlos De Martin
ICASSP2
2002 Delivery of MPEG video streams with constant perceptual quality of service
abstract
Constant levels of perceptual quality of service is what ideally users of multimedia services expect. In most cases, however, they receive time-varying levels of quality of service. This paper describes a technique to deliver nearly constant perceptual quality of service when transmitting video sequences over differentiated services IP networks. MPEG video packets are transmitted either as low-loss premium packets or as regular best-effort packets depending on their individual perceptual importance. On a frame-by-frame basis, allocation to the premium class is performed depending on the perceptual importance of each macroblock, the desired level of quality of service and the instantaneous network state. The resulting perceptually-based, time-varying use of premium and best-effort network resources delivers nearly constant quality of service to end users and it yields significant higher PSNR values compared to constant allocation of premium bandwidth.
Davide Quaglia, Juan Carlos De Martin
ICME (2)2
2002 Perceptual classification of MPEG video for Differentiated-Services communications
abstract
We present a distortion-based packet marking technique for transmission of motion-compensated video over Differentiated Services networks. For each macroblock of an MPEG2 video sequence, the distortion that would be caused at the receiver by its loss is computed. High distortion macroblocks are grouped into perceptually important slices that can be transmitted as premium packets, while lower distortion slices are sent as less expensive, best-effort traffic. Firstly, computation of the distortion introduced in the current frame only is compared to exhaustive computation of the distortion introduced in the entire group of pictures (GOP) due to the error propagation. Secondly, allocation of the premium traffic on a frame-by-frame basis is compared to GOP-wide allocation. Results show that GOP-wide allocation of premium traffic is key in using premium bandwidth efficiently, with strong PSNR gains with respect to the other approaches. We also propose a model-based distortion computation technique, which, combined with GOP-level premium traffic allocation, delivers nearly the same performance of the exhaustive approach at a fraction of its complexity.
Fabio De Vito, Laura Farinetti, Juan Carlos De Martin
ICME (1)3
2002 Perception-based partial encryption of compressed speech
abstract
Mobile multimedia applications, the focus of many forthcoming wireless services, increasingly demand low-power techniques implementing content protection and customer privacy. In this paper low complexity perception-based partial encryption schemes for speech are presented. Speech compressed by a widely-used speech coding algorithm, the ITU-T G.729 standard at 8 kb/s, is partitioned in two classes, one, the most perceptually relevant, to be encrypted, the other, to be left unprotected. Two partial-encryption techniques are developed, a low-protection scheme, aimed at preventing most kinds of eavesdropping and a high-protection scheme, based on the encryption of a larger share of perceptually important bits and meant to perform as well as full encryption of the compressed bitstream. The high-protection scheme, based on the encryption of about 45% of the bitstream, achieves content protection comparable to that obtained by full encryption, as verified by both objective measures and formal listening tests. For the low-protection scheme, encryption of as little as 30% of the bitstream virtually eliminates intelligibility as well as most of the remaining perceptual information. Low-power, portable devices could therefore achieve very high levels of speech-content protection at only 30-45% of the computational load of current techniques, freeing resources for other tasks and enabling longer battery life.
Antonio Servetti, Juan Carlos De Martin
IEEE Trans. Speech Audio Process.2
2001 Source-driven packet marking for speech transmission over differentiated-services networks
abstract
We present a source-driven approach to packet marking for speech transmission over packet networks implementing the Differentiated Services model. Packets generated by the speech coder are examined: if deemed perceptually critical, they are marked as premium and sent on a "virtual wire;" otherwise, they are sent as regular best-effort traffic. Applied to speech coded with the ITU-T 8 kb/s speech coding standard G.729, the proposed source-driven packet marking scheme outperforms source-transparent techniques and provides clearly better perceptual quality than the unprotected case sending as little as 1/5 of the coded bitstream as premium traffic.
Juan Carlos De Martin
ICASSP1
2001 Distortion-Based Packet Marking For Mpeg Video Transmission Over Diffserv Networks
abstract
We present a distortion-based approach to packet classification for multimedia transmission over differentiatedservices packet networks. Instead of sending all traffic as premium or relying on a priori data partitioning, packets are individually examined and assigned to different service classes depending on the level of distortion that their loss would introduce at the decoder. Applied to video sequences encoded with the ISO MPEG-2 video coding standard, the proposed distortion-based packet marking scheme outperforms source-transparent techniques and provides substantial and consistent gains in PSNR over the regular best-effort case sending as little as 10% of the packets as premium traffic. Video samples are available at #####################################.
Juan Carlos De Martin, Davide Quaglia
ICME1
2001 Adaptive picture slicing for distortion-based classification of video packets
abstract
We present a new algorithm for the dynamic identification of perceptually important regions in pictures of a video sequence. For each macroblock, the distortion that its loss would cause at the decoder is computed. High-distortion macroblocks are then grouped together into slices to be protected with forward error correction or to be sent as "premium" packets. We developed the approach for the case of video sequences encoded with the ISO MPEG-2 video coding standard. We then applied the algorithm to classify video packets within a 1-bit differentiated services architecture: slices were grouped either into premium packets, to be sent on a "virtual wire", or into regular packets, to be sent as best-effort traffic. In packet losses, the proposed distortion-based classification scheme outperforms source-transparent packet-marking techniques and provides substantially higher PSNR values than the regular best-effort case sending as little as 10% of the packets as premium traffic. Video samples are available at http://multimedia.polito.it/mmsp2001/.
Enrico Masala, Davide Quaglia, Juan Carlos De Martin
MMSP3
2001 A simulation study of adaptive voice communications on IP networks
A. Barberis, Claudio Casetti, Juan Carlos De Martin, Michela Meo
Comput. Commun.3
2000 Improved frame erasure concealment for CELP-based coders
abstract
This paper describes new techniques for concealing frame erasures for CELP-based speech coders. Two main approaches were followed: interpolative, where both past and future information are used to reconstruct the missing data, and repetition-based, where no future information is required. Key features of the repetition-based approach include improved muting, pitch delay jittering, and LPC bandwidth expansion. The interpolative approach can be employed in voice over IP scenarios at no extra cost in terms of delay. Applied to the ITU-T G.729 ACELP 8 kb/s speech coding standard, both interpolation- and repetition-based techniques outperform standard concealment in informal listening tests.
Juan Carlos De Martin, Takahiro Unno, Vishu Viswanathan
ICASSP1
2000 A Framework for the Analysis of Adaptive Voice over IP
abstract
We present a framework for the analysis of a set of adaptive variable-bit-rate voice sources in a packet network. The instantaneous bit rate of each source is determined by an end-to-end control mechanism that, based on measurements of packet delay and loss rate, selects the rate that best matches current network conditions. Several such algorithms can be analyzed with the proposed framework, which consists of a detailed Markovian model of the source behavior and of an approximate description of the interaction between the sources and the underlying network. The model of the source takes into account time intervals during which a connection is active as well as intervals of inactivity; within a given conversation, it also models on/off (speech/silence) periods. The interaction of a source with the rest of the system is derived through an iterative procedure that evaluates the feedback that a source receives from the network. A case study presenting the results relative to an adaptive system transmitting at bit rates typical of widely used speech coding standards (64 kb/s, 13 kb/s and 8 kb/s) illustrates the proposed framework.
Claudio Casetti, Juan Carlos De Martin, Michela Meo
ICC (2)2
1999 An adaptive multi-rate speech coder for digital cellular telephony
abstract
We have developed an adaptive multi-rate (AMR) speech coder designed to operate under the GSM digital cellular full rate (22.8 kb/s) and half rate (11.4 kb/s) channels and to maintain high quality in the presence of highly varying background noise and channel conditions. Within each total rate, several codec modes with different source/channel bit rate allocations are used. The speech coders in each codec mode are based on the CELP algorithm operating at rates ranging from 11.85 kb/s down to 5.15 kb/s, where the lowest rate coder is a source controlled multi-modal speech coder. The decoders monitor the channel quality at both ends of the wireless link using the soft values for the received bits and assist the base station in selecting the codec mode that is appropriate for a given channel condition. The coder was submitted to the GSM AMR standardization competition and met the qualification requirements in an independent formal MOS test.
Erdal Paksoy, Juan Carlos De Martin, Alan McCree, Christian G. Gerlach, Anand K. Anandakumar, Wai-Ming Lai, Vishu Viswanathan
ICASSP2
1998 A 1.7 kb/s MELP coder with improved analysis and quantization
abstract
This paper describes our new mixed excitation linear predictive (MELP) coder designed for very low bit rate applications. This new coder, through algorithmic improvements and enhanced quantization techniques, produces better speech quality at 1.7 kb/s than the new U.S. Federal Standard MELP coder at 2.4 kb/s. Key features of the coder are an improved pitch estimation algorithm and a line spectral frequencies (LSF) quantization scheme that requires only 21 bits per frame. With channel coding, this new MELP coder is capable of maintaining good speech quality even in severely degraded channels, at a total bit rate of only 3 kb/s.
Alan McCree, Juan Carlos De Martin
ICASSP2
1998 A mixed sinusoidally excited linear prediction coder at 4 kb/s and below
abstract
There is currently a great deal of interest in the development of speech coding algorithms capable of delivering toll quality at 4 kb/s and below. For synthesizing high quality speech, accurate representation of the voiced portions of speech is essential. For bit rates of 4 kb/s and below, conventional code excited linear prediction (CELP) may likely not provide the appropriate degree of periodicity. It has been shown that good quality low bit rate speech coding can be obtained by frequency domain techniques such as sinusoidal transform coding (STC), multi-band excitation (MBE), mixed excitation linear prediction (MELP), and multi-band LPC (MB-LPC) vocoders. In this paper, a speech coding algorithm based on an improved version of MB-LPC is presented. Main features of this algorithm include a multi-stage time/frequency pitch estimation and an improved mixed voicing representation. An efficient quantization scheme for the spectral amplitudes of the excitation, called formant weighted vector quantization, is also used. This improved coder, called mixed sinusoidally excited linear prediction (MSELP), yields an unquantized model with speech quality better than the 32 kb/s AD-PCM quality. Initial efforts towards a fully quantized 4 kb/s coder, although not yet successful in achieving the toll quality goal, have produced good output speech quality.
Suat Yeldener, Juan Carlos De Martin, Vishu Viswanathan
ICASSP2
1996 Mixed-domain coding of speech at 3 kb/s
abstract
We present a speech coding algorithm called mixed-domain residual coding (MDRC) wherein a prototype pitch cycle in each frame of the speech residual is coded in the time-domain while interpolation of the residual signal is performed in the frequency-domain. A novel quantization scheme takes into account time scaling and differentially codes successive prototypes with a closed-loop perceptually-weighted search. A fixed-rate (3.15 kb/s) implementation of MDRC achieves quality better or comparable to higher rate coders such as FS 1016 CELP and IMBE.
Juan Carlos De Martin, Allen Gersho
ICASSP1