EDBT 2026 Demo / reviewers in the wild / expert
Emre Aksu
dblp:192/1480 · also Emre B. Aksu, Emre Baris Aksu
· DBLP profile ↗
26ranked-venue papers
0as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 15 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Metadata-Guided Hot Swapping of Specialized Super-Resolution Models in Streaming SystemsabstractStreaming systems that employ video super-resolution (SR) often rely on a single, generic neural network model for all content types, resulting in suboptimal visual quality across diverse scenes. To address this limitation, we propose a metadata-guided hot swapping mechanism that enables the dynamic selection of specialized, fine-tuned SR models during streaming. The system uses content type change signals transmitted via an auxiliary metadata track, which is prioritized over media tracks to ensure early arrival. This allows the client to preload the appropriate neural network before the content type changes, minimizing startup delay and improving responsiveness. The end-to-end workflow includes content detection through source mapping, high-priority metadata transmission, message extraction, neural network preloading and SR application. Alperen F. Zengin, Ekrem Çetinkaya, Ali C. Begen, Saba Ahsan, Serhan Gul, Kashyap Kammachi Sreedhar, Emre Aksu |
ISM | 7 |
| 2023 | Quality Upshifting with Auxiliary I-Frame SplicingabstractThis paper introduces the Auxiliary I-Frame Splicing method to reduce bandwidth waste in adaptive streaming. This method involves fetching a high-quality I-frame and splicing it into the already downloaded low-quality segment, resulting in a higher-quality rendering at a lower overhead than replacing the entire low-quality segment. In our experiments with three videos and four quantization parameters, the results show that the bandwidth can be saved up to 87% while still increasing the peak signal-to-noise ratio score by 20% and the video multi-method assessment fusion score by 73%. In the demo, we demonstrate the visual differences between the original and spliced videos. Mehmet N. Akcay, Burak Kara, Ali C. Begen, Saba Ahsan, Igor D. D. Curcio, Kashyap Kammachi Sreedhar, Emre Aksu |
QoMEX | 7 |
| 2022 | Bridging the Gap Between Image Coding for Machines and HumansabstractImage coding for machines (ICM) aims at reducing the bitrate required to represent an image while minimizing the drop in machine vision analysis accuracy. In many use cases, such as surveillance, it is also important that the visual quality is not drastically deteriorated by the compression process. Recent works on using neural network (NN) based ICM codecs have shown significant coding gains against traditional methods; however, the decompressed images, especially at low bitrates, often contain checkerboard artifacts. We propose an effective decoder finetuning scheme based on adversarial training to significantly enhance the visual quality of ICM codecs, while preserving the machine analysis accuracy, without adding extra bitcost or parameters at the inference phase. The results show complete removal of the checkerboard artifacts at the negligible cost of −1.6% relative change in task performance score. In the cases where some amount of artifacts is tolerable, such as when machine consumption is the primary target, this technique can enhance both pixel-fidelity and feature-fidelity scores without losing task performance. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ICIP | 6 |
| 2022 | Stochastic Binary-Ternary Quantization for Communication Efficient Federated ComputationabstractA stochastic binary-ternary (SBT) quantization approach is introduced for communication efficient federated computation; form of collaborative computing where locally trained models are exchanged between institutes. Communication of deep neural network models could be highly inefficient due to their large size. This motivates model compression in which quantization is an important step. Two well-known quantization algorithms are binary and ternary quantization. The first leads into good compression, sacrificing accuracy. The second provides good accuracy with less compression. To better benefit from trade-off between accuracy and compression, we propose an algorithm to stochastically switch between binary and ternary quantization. By combining with uniform quantization, we further extend the proposed algorithm to a hierarchical method which results in even better compression without sacrificing the accuracy. We tested the proposed algorithm using Neural network Compression Test Model (NCTM) provided by MPEG community. Our results demonstrate that the hierarchical variant of the proposed algorithm outperforms other quantization algorithms in term of compression, while maintaining the accuracy competitive to that provided by other methods. Rangu Goutham, Homayun Afrabandpey, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Miska M. Hannuksela, Hamed Rezazadegan Tavakoli |
ICIP | 5 |
| 2022 | TMD: Transformed Mesh Decoder for Mesh AnimationabstractEasy and fast animation of 3D characters is attractive for both gaming and entertainment applications. 3D mesh must be properly rigged and skinned in order to create seamless animation. This process can be time consuming and requires deep knowledge of appropriate software based on kinematic animation. In this work, we present a fast and lightweight deep neural model to automate 3D human animation using skeletal representation from 2D image pose, i.e., joints in 2D space. We accomplish this using Transformed Mesh Decoder (TMD), which is a novel layer for convolutional neural networks. To train the network, we generate a large and diverse dataset using Skinned Multi-Person Linear (SMPL) model. Experiment shows that our method is effective when compared to both the ground truth and state-of-the art linear blend skinning that require manually painted skinning weights for accurate result. The animation process is fast and can achieve approximately 10-15fps in practice. The proposed method is simple which opens the possibility for future improvement in real-time application. Peter Fasogbon, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu |
ICPR | 5 |
| 2022 | Benchmarking the Second Edition of the Omnidirectional Media Format StandardabstractOmnidirectional MediA Format (OMAF) is the first worldwide virtual reality (VR) standard to store and distribute immersive media, completed in 2019. Later, in 2021, the second edition of this standard (OMAF v2) was published. The second edition kept all the features defined in the first OMAF edition while introducing some new ones, such as overlays and multi-viewpoints. OMAF v2’s Tile Index Segments that contain metadata to track fragment data per segment and quality levels create a bandwidth overhead. During the OMAF v2 standardization, multiple methods for the track fragment run representation were studied to deal with this overhead. This paper presents the implementation of one of these methods, the compressed box method using the DEFLATE algorithm (OMAF v2*). It also provides comprehensive test results of OMAF v1, OMAF v2 and OMAF v2* with various combinations of three tile grids (6x4, 8x6 and 12x8), three segment durations (300 ms, 900 ms and 3 s), two videos (RollerCoaster and Timelapse), two bitrate groups (each group with four different bitrates) and two HTTP versions (HTTP/1.1 and H2). Burak Kara, Mehmet N. Akcay, Ali C. Begen, Saba Ahsan, Igor D. D. Curcio, Kashyap Kammachi Sreedhar, Emre Aksu |
ISM | 7 |
| 2022 | Optimizing storage and delivery of Omnidirectional Videos in Viewport-dependent streamingabstractThe OMAF standard makes use of a framework called the viewport-dependent-delivery for the streaming of 360-degree videos. OMAF uses ISOBMFF for storage and MPEG-DASH as one of the delivery mechanisms. In viewport-dependent-streaming videos are spatially divided and encoded into multiple tracks and each track is further segmented for DASH delivery. Segmentation requires additional metadata which adds to bitrate overhead. The main contributor to this overhead is the track fragment run in a box with the four-character code, ‘trun’. The TRUN records the following information of each sample in a track: the size, duration, flags, and time offsets and uses a fixed byte size to record this information. To minimize the bitrate overhead of TRUN, four different representation algorithms have been explored. This paper briefly describes the four TRUN representations and discusses the benefits and drawbacks of each algorithm. For evaluation, the algorithms were implemented in the MP4BOX module of the GPAC suite. The results were evaluated for different segment durations (500ms, 1s, 2s, 4s), different tiling grids (8x4, 9x6), two videos (bip-bop, countertiles) with different packaging techniques (no encryption, encryption of Keyframes, encryption of all frames) The algorithms reduced the bitrate overhead by 59% on average as compared to the original TRUN representation. Kashyap Kammachi Sreedhar, Miska M. Hannuksela, Emre Aksu, Lauri Ilola, Lukasz Condrad |
ISM | 3 |
| 2022 | On the Importance of Temporal Dependencies of Weight Updates in Communication Efficient Federated LearningabstractThis paper studies the effect of exploiting temporal dependency of successive weight updates on compressing communications in Federated Learning (FL). For this, we propose residual coding for FL, which utilizes temporal dependencies by communicating compressed residuals of the weight updates whenever they are beneficial to bandwidth. We further consider Temporal Context Adaptation (TCA) which compares co-located elements of consecutive weight updates to select optimal setting for compression of bitstream in DeepCABAC encoder. Following experimental settings of MPEG standard on Neural Network Compression (NNC), we demonstrate that both temporal dependency based technologies reduce communication overhead, where the maximum reduction is obtained using both technologies, simultaneously. Homayun Afrabandpey, Rangu Goutham, Honglei Zhang 0001, Francesco Criri, Emre Aksu, Hamed Rezazadegan Tavakoli |
VCIP | 5 |
| 2022 | Overview of the Neural Network Compression and Representation (NNR) StandardabstractNeural Network Coding and Representation (NNR) is the first international standard for efficient compression of neural networks (NNs). The standard is designed as a toolbox of compression methods, which can be used to create coding pipelines. It can be either used as an independent coding framework (with its own bitstream format) or together with external neural network formats and frameworks. For providing the highest degree of flexibility, the network compression methods operate per parameter tensor in order to always ensure proper decoding, even if no structure information is provided. The NNR standard contains compression-efficient quantization and deep context-adaptive binary arithmetic coding (DeepCABAC) as core encoding and decoding technologies, as well as neural network parameter pre-processing methods like sparsification, pruning, low-rank decomposition, unification, local scaling and batch norm folding. NNR achieves a compression efficiency of more than 97% for transparent coding cases, i.e. without degrading classification quality, such as top-1 or top-5 accuracies. This paper provides an overview of the technical features and characteristics of NNR. Heiner Kirchhoffer, Paul Haase, Wojciech Samek, Karsten Müller 0001, Hamed Rezazadegan Tavakoli, Francesco Cricri, Emre Aksu, Miska M. Hannuksela, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001, Swayambhoo Jain, Shahab Hamidi-Rad, Fabien Racapé, Werner Bailer |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2021 | Mind The Structure: Adopting Structural Information For Deep Neural Network CompressionabstractDeep neural networks have huge number of parameters and require large number of bits for representation. This hinders their adoption in decentralized environments where model transfer among different parties is a characteristic of the environment while the communication bandwidth is limited. Parameter quantization is a compression approach to address this challenge by reducing the number of bits required to represent a model, e.g. a neural network. However, majority of existing neural network quantization methods do not exploit structural information of layers and parameters during quantization. In this paper, focusing on Convolutional Neural Networks (CNNs), we present a novel quantization approach by employing the structural information of neural network layers and their corresponding parameters. Starting from a pre-trained CNN, we categorize network parameters into different groups based on the similarity of their layers and their spatial structure. Parameters of each group are independently clustered and the centroid of each cluster is used as representative for all parameters in the cluster. Finally, the centroids and the cluster indexes of the parameters are used as a compact representation of the parameters. Experiments with two different tasks, i.e., acoustic scene classification and image compression, demonstrate the effectiveness of the proposed approach. Homayun Afrabandpey, Anton Muravev, Hamed Rezazadegan Tavakoli, Honglei Zhang 0001, Francesco Cricri, Moncef Gabbouj, Emre Aksu |
ICIP | 7 |
| 2021 | Hybrid Pruning And SparsificationabstractA hybrid approach based on the combination of saliency-based neural pruning and regularization-based sparsification is proposed. We propose using a graph diffusion process for determining the neuron importance for pruning. Then, we use a regularization loss based on weighted $L_{1}-$norm and $L_{2}-$norm during fine-tuning to recover the lost performance. This is followed by a threshold step to further impose sparsification. We demonstrate such a hybrid approach achieves significantly better performance in comparison to purely regularization-based sparsification for large neural networks. To this end, we assessed our proposed method on three tasks, including: image classification (3 network architectures), audio classification and image compression. Hamed Rezazadegan Tavakoli, Joachim Wabnig, Francesco Cricri, Honglei Zhang 0001, Emre Aksu, Iraj Saniee |
ICIP | 5 |
| 2021 | Enhancing Image Coding for Machines with Compressed Feature ResidualsabstractAs computer vision technologies have tremendously improved over the last decade, videos and images are often consumed by machines instead of humans which are the main target for traditional video codecs. In many use cases, although machines are the main consumers, human involvement is also required, or even mandatory. In this paper, we propose a novel image coding technique targeted for machines, while maintaining the capability for human consumption. Our proposed codec generates two bitstreams: one bitstream from a traditional codec, referred to as human bitstream, optimized for human consumption; the other bitstream, referred to as machine bitstream, generated from an end-to-end learned neural network-based codec and optimized for machine tasks. Instead of working on the image domain, the proposed machine bitstream is derived from feature residuals – the difference between the features extracted from the input image and the features extracted from the reconstructed image generated by the traditional codec. With the help of the machine bitstream, we can significantly improve machine task performance in the low bitrate range. Our system beats the state-of-the-art traditional codec, the Versatile Video Coding (VVC/H.266), achieving −40.5% in Bjontegaard delta bitrate reduction on average for bitrates up to 0.07 BPP. Joni Seppälä, Honglei Zhang 0001, Nam Le 0003, Ramin Ghaznavi Youvalari, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 7 |
| 2021 | Adaptation and Attention for Neural Video CodingabstractNeural image coding represents now the state-of-the-art image compression approach. However, a lot of work is still to be done in the video domain. In this work, we propose an end-to-end learned video codec that introduces several architectural novelties as well as training novelties, revolving around the concepts of adaptation and attention. Our codec is organized as an intra-frame codec paired with an inter-frame codec. As one architectural novelty, we propose to train the inter-frame codec model to adapt the motion estimation process based on the resolution of the input video. A second architectural novelty is a new neural block that combines concepts from split-attention based neural networks and from DenseNets. Finally, we propose to overfit a set of decoder-side multiplicative parameters at inference time. Through ablation studies and comparisons to prior art, we show the benefits of our proposed techniques in terms of coding gains. We compare our codec to VVC/H.266 and RLVC, which represent the state-of-the-art traditional and end-to-end learned codecs, respectively, and to the top performing end-to-end learned approach in 2021 CLIC competition, E2E_T_OL. Our codec clearly outperforms E2E_T_OL, and compare favorably to VVC and RLVC in some settings. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Jani Lainema, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 7 |
| 2021 | Head-Motion-Aware Viewport Margins for Improving User Experience in Immersive VideoabstractViewport-dependent delivery (VDD) is a technique to save network resources during the transmission of immersive videos. However, it results in a non-zero motion-to-high-quality delay (MTHQD), which is the delta time from the moment where the current viewport has at least one low-quality tile to when all the tiles in the new viewport are rendered in high quality. MTHQD is an important metric in the evaluation of the VDD systems. This paper improves an earlier concept called viewport margins by introducing head-motion awareness. The primary benefit of this improvement is the reduction (up to 64%) in the average MTHQD. Mehmet N. Akcay, Burak Kara, Saba Ahsan, Ali C. Begen, Igor D. D. Curcio, Emre Aksu |
MMAsia | 6 |
| 2021 | NBMP Standard Use Case: 3D Human Reconstruction WorkflowabstractWe present a demonstration of Network Based Media Processing (NBMP) standard-compliant cloud service for reconstructing and Augmented Reality (AR) display of fully textured 3D human model using 2-5 images captured with a smartphone. Yu You, Peter Fasogbon, Emre Aksu |
VCIP | 3 |
| 2020 | Lossless Image Compression Using a Multi-scale Progressive Statistical Model
Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Nannan Zou, Emre Aksu, Miska M. Hannuksela |
ACCV (3) | 5 |
| 2020 | L2C - Learning to Learn to CompressabstractIn this paper we present an end-to-end meta-learned system for image compression. Traditional machine learning based approaches to image compression train one or more neural network for generalization performance. However, at inference time, the encoder or the latent tensor output by the encoder can be optimized for each test image. This optimization can be regarded as a form of adaptation or benevolent overfitting to the input content. In order to reduce the gap between training and inference conditions, we propose a new training paradigm for learned image compression, which is based on meta-learning. In a first phase, the neural networks are trained normally. In a second phase, the Model-Agnostic Meta-learning approach is adapted to the specific case of image compression, where the inner-loop performs latent tensor overfitting, and the outer loop updates both encoder and decoder neural networks based on the overfitting performance. Furthermore, after meta-learning, we propose to overfit and cluster the bias terms of the decoder on training image patches, so that at inference time the optimal content-specific bias terms can be selected at encoder-side. Finally, we propose a new probability model for lossless compression, which combines concepts from both multi-scale and super-resolution probability model approaches. We show the benefits of all our proposed ideas via carefully designed experiments. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Jani Lainema, Miska M. Hannuksela, Emre Aksu, Esa Rahtu |
MMSP | 7 |
| 2019 | Demo: Accelerating Depth-Map on Mobile Device Using CPU-GPU Co-processing
Peter Fasogbon, Emre Aksu, Lasse Heikkilä |
CAIP (1) | 2 |
| 2019 | Calibration of Fisheye Camera Using Entrance PupilabstractMost conventional camera calibration algorithms assume that the imaging device has a Single Viewpoint (SVP). This is not necessarily true for special imaging device such as fisheye lenses. As a consequence, the intrinsic camera calibration result is not always reliable. In this paper, we propose a new formation model that tends to relax this assumption so that a Non-Single Viewpoint (NSVP) system is corrected to always maintain a SVP, by taking into account the variation of the Entrance Pupil (EP) using thin lens modeling. In addition, we present a calibration procedure for the image formation to estimate these EP parameters using non linear optimization procedure with bundle adjustment. From experiments, we are able to obtain slightly better re-projection error than traditional methods, and the camera parameters are better estimated. The proposed calibration procedure is simple and can easily be integrated to any other thin lens image formation model. Peter Fasogbon, Emre Aksu |
ICIP | 2 |
| 2019 | Portrait Instance Segmentation for Mobile DevicesabstractAccurate and efficient portrait instance segmentation has become a crucial enabler for many multimedia applications on mobile devices. We present a novel convolutional neural network (CNN) architecture to explicitly address the long standing problems in portrait segmentation, i.e., semantic coherence and boundary localization. Specifically, we propose a cross-granularity categorical attention mechanism leveraging the deep supervisions to close the semantic gap of CNN feature hierarchy by imposing consistent category-oriented information across layers. Furthermore, a cross-granularity boundary enhancement module is proposed to boost the boundary awareness of deep layers by integrating the shape context cues from shallow layers of the network. We further propose a novel and efficient non-parametric affinity model to achieve efficient instance segmentation on mobile devices. We present a portrait image dataset with instance level annotations dedicated to evaluating portrait instance segmentation algorithms. We evaluate our approach on challenging datasets which obtains state-of-the-art results. Lingyu Zhu 0001, Tinghuai Wang, Emre Aksu, Joni-Kristian Kämäräinen |
ICME | 3 |
| 2019 | Frame selection to accelerate Depth from Small Motion on smartphonesabstractDepth from Small Motion (DfSM) is particularly interesting for smartphone devices because it makes it possible to get depth information with minimal user effort and cooperation. The state of art method requires about 30 images for the optimization to converge fast and produce accurate depth-map. As the use of high number of frames contribute to long execution time and huge memory allocation, we propose a frame selection strategy using Inertial Measurement Unit (IMU) and image based analysis. As a result, only 5 frames with appropriate viewpoint from the reference one are used for the depth-map generation. Full experiment is done on an Android platform using optimized version of the proposed method with CPU-GPU co-processing under OpenCL. We are able to provide accurate camera parameters and depth-map estimates using only these selected frames. Peter Fasogbon, Lasse Heikkilä, Emre Aksu |
IECON | 3 |
| 2019 | Approximating Binarization in Neural NetworksabstractBinarization of neural networks' activations may be a requirement for some applications. A typical example is end-to-end learned deep image compression systems where the encoder's output is requred to be a binary vector. Binarization is non-differentiable, therefore one needs to approximate it in order to train neural networks with stochastic gradient descent. In this paper, we investigate these training strategies and provide improvements over baselines. We find that during training, constraining the activations in a region that is far away from binary points leads to a better performance at test-time. The above finding provides a counter-intuitive result and leads to re-thinking the binarization approximation problem in neural networks. Çaglar Aytekin, Francesco Cricri, Jani Lainema, Emre Aksu, Miska M. Hannuksela |
IJCNN | 4 |
| 2018 | Clustering and Unsupervised Anomaly Detection with l2 Normalized Deep Auto-Encoder RepresentationsabstractClustering is essential to many tasks in pattern recognition and computer vision. With the advent of deep learning, there is an increasing interest in learning deep unsupervised representations for clustering analysis. Many works on this domain rely on variants of auto-encoders and use the encoder outputs as representations/features for clustering. In this paper, we show that an l2normalization constraint on these representations during auto-encoder training, makes the representations more separable and compact in the Euclidean space after training. This greatly improves the clustering accuracy when k-means clustering is employed on the representations. We also propose a clustering based unsupervised anomaly detection method using l2normalized deep auto-encoder representations. We show the effect of l2normalization on anomaly detection accuracy. We further show that the proposed anomaly detection method greatly improves accuracy compared to previously proposed deep methods such as reconstruction error based anomaly detection. Çaglar Aytekin, Xingyang Ni, Francesco Cricri, Emre Aksu |
IJCNN | 4 |
| 2018 | Memory-Efficient Deep Salient Object Segmentation Networks on Gridized SuperpixelsabstractComputer vision algorithms with pixel-wise labeling tasks, such as semantic segmentation and salient object detection, have gone through a significant accuracy increase with the incorporation of deep learning. Deep segmentation methods slightly modify and fine-tune pre-trained networks that have hundreds of millions of parameters. In this work, we question the need of having such memory demanding networks for a reasonable performance in salient object segmentation. To this end, we propose a way to learn a memory-efficient network from scratch by training it only on salient object detection datasets. Our method encodes images to gridized superpixels that preserve both the object boundaries and the connectivity rules of regular pixels. This representation allows us to use convolutional neural networks that operate on regular grids. By using these encoded images, we train a memory-efficient network using only 0.048% of the number of parameters that a majority of other deep salient object detection networks have. Our method shows comparable accuracy with the state-of-the-art deep salient object detection methods and provides a much more memory-efficient alternative to them. Due to its easy deployment and small size, such a network is preferable for applications in memory limited IoT devices. Çaglar Aytekin, Xingyang Ni, Francesco Cricri, Lixin Fan, Emre Aksu |
MMSP | 5 |
| 2016 | HEVC still image coding and high efficiency image file formatabstractThe High Efficiency Video Coding (HEVC) standard includes support for a large range of image representation formats and provides an excellent image compression capability. The High Efficiency Image File Format (HEIF) offers a convenient way to encapsulate HEVC coded images, image sequences and animations together with associated metadata into a single file. This paper discusses various features and functionalities of the HEIF file format and compares the compression efficiency of HEVC still image coding to that of JPEG 2000. According to the experimental results HEVC provides about 25% bitrate reduction compared to JPEG 2000, while keeping the same objective picture quality. Jani Lainema, Miska M. Hannuksela, Vinod Kumar Malamal Vadakital, Emre Aksu |
ICIP | 4 |
| 1999 | Implementation of an enhanced fixed point variable bit-rate MELP vocoder on TMS320C549abstractIn this paper, a fixed point variable bit-rate (VBR) mixed excitation linear predictive coding (MELP/sup TM/) vocoder is presented. The VBR-MELP vocoder is also implemented on a TMS320C54x and it achieves virtually indistinguishable federal standard MELP quality at bit-rates between 1.0 to 1.6 kb/s. The backbone of VBR-MELP vocoder is similar to that of federal standard MELP. It utilizes a novel sub-band based voice activity detector in the back-end of encoder to discriminate background noise from speech activity. Since proposed detector uses only parameters extracted in the encoder, its computational complexity is very low. Ali Erdem Ertan, Emre Aksu, Hakki Gökhan Ilk, Mehmet Haydar Karcí, Onder Karpat, Taner Kolcak, Levent Sendur, Mübeccel Demirekler, A. Enis Çetin |
ICASSP | 2 |