EDBT 2026 Demo / reviewers in the wild / expert
Moncef Gabbouj
dblp:08/6597
· DBLP profile ↗
393ranked-venue papers
6as first author
91since 2021 · last 2026
0000-0002-9788-2323ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 256 · 5 first-author · 44 since 2021Artificial intelligence and machine learning · 108 · 1 first-author · 34 since 2021Systems, architecture and hardware · 22 · 3 since 2021Computer networks · 13 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Security and privacy · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JNTD-DS: A Benchmark Dataset for Just Noticeable frame rate-based Temporal Difference in Perceptual Video CodingabstractThe Just Noticeable Difference (JND) is defined as the maximum change in a visual stimulus (image or video) which the Human Visual System (HVS) can tolerate without perceiving visual distortion. Previous JND research has largely focused on spatial distortions, yielding several datasets and models that predict spatial thresholds based on parameters such as Quantization Parameter (QP) or Quality Factor (QF). However, temporal thresholds, specifically the maximum frame rate reductions which viewers cannot detect, remain largely unexplored, despite their critical importance for efficient video coding. To address this gap, we introduce JNTD-DS, which, to the best of our knowledge, is the first benchmark dataset specifically designed to measure the Just Noticeable frame rate-based Temporal Difference (JNTD). The dataset comprises 50 video scenes covering various content, and the JNTD level associated with them. T e video scenes are studied through extensive subjective tests, comparing the high frame rate videos with their temporally downsampled versions. This forms 1196 opinion scores from 78 subjects. Analyzing the collected data confirms that JNTD thresholds, which are fundamentally defined by the HVS, are inherently complex and vary across content. By providing critical insights into HVS sensitivity to frame rate changes, the dataset enables content-adaptive frame rate optimization for perceptual video coding, allowing more efficient compression in video streaming and bandwidth-limited applications without compromising visual quality. We further demonstrate the practical impact of these insights by developing a JNTD prediction model and integrating it into a video compression pipeline, achieving an average bitrate reduction of 13.62% with only a marginal quality loss. The JNTD-DS is publicly available at https://github.com/sanaznami/JNTD-DS. Sanaz Nami, Farhad Pakdaman, Sahab Taali, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
MMSys | 6 |
| 2026 | Comparative analysis of BLE SIG mesh and Wirepas mesh for Ad Hoc IoT deployments: Features, security, efficiency and application suitabilityabstractThe rapid increase in the number of Internet of Things (IoT) devices has led to the development of advanced networking technologies such as bluetooth low energy (BLE) standardized by special interest group (SIG) mesh and Wirepas mesh networks. Each of these technologies offers unique features and capabilities. BLE is a widely used short-range technology that has made a significant impact on the IoT paradigm development thanks to its simplicity, low power consumption, robustness, and low cost. In contrast, Wirepas mesh is a decentralized wireless communication protocol, meaning that each node in the network selects its own role while maintaining and optimizing the connection automatically based on its environment. This paper provides a comparative analysis of BLE SIG mesh and Wirepas mesh, focusing on their security, scalability, power, and memory efficiency, as well as application suitability in diverse ad hoc IoT environments. The study highlights that while BLE SIG mesh benefits from easy adoption, low power consumption, and ease of integration into the consumer IoT ecosystem, Wirepas mesh excels in high-density, large-scale industrial applications due to its robust management and decentralized control. Security is realized in BLE SIG mesh with standard authentication and encryption, whereas Wirepas mesh utilizes deployment-configurable security mechanisms to handle diverse network topologies. This comparative review enables IoT industrialists and researchers to select appropriate mesh technologies based on their specific requirements and deployment constraints. Muhammad Zeeshan Waheed, Fahad Sohrab, Waleed Bin Qaim, Matti Vakkuri, Jyrki Okkonen, Mikko Valkama, Moncef Gabbouj |
Ad Hoc Networks | 7 |
| 2026 | FANeRV: frequency separation and augmentation based neural representation for video
Li Yu 0004, Jimin Xiao, Moncef Gabbouj |
Expert Syst. Appl. | 5 |
| 2026 | Neural architecture search for generative adversarial networks with hybrid convolution
Yu Xue 0003, Yufeng Zou, Mohamed Wahib, Peng Chen 0035, Moncef Gabbouj |
Neurocomputing | 5 |
| 2026 | Deep learning-based point cloud upsampling: A survey of methodologies, performance comparisons, and noise robustness analysis
Yihang Yin, Li Yu 0004, Wei Zhou 0021, Moncef Gabbouj |
Neurocomputing | 4 |
| 2026 | DRACO: Data Replication and Collection Framework for Enhanced Data Availability and Robustness in IoT NetworksabstractThe Internet of Things (IoT) bridges the gap between the physical and digital worlds, enabling seamless interaction with real-world objects via the Internet. However, IoT systems face significant challenges in ensuring efficient data generation, collection, and management, particularly due to the resource-constrained and unreliable nature of connected devices, which can lead to data loss. This paper presents DRACO (Data Replication and Collection), a framework that integrates a distributed hop-by-hop data replication approach with a routing-free mobile sink-based data collection strategy. DRACO enhances data availability, optimizes replica placement, and ensures efficient data retrieval even under node failures and varying network densities. Extensive ns-3 simulations demonstrate that DRACO outperforms state-of-the-art techniques, improving data availability by up to 15% and 34%, and replica creation by up to 18% and 40%, compared to greedy and random replication techniques, respectively. DRACO also ensures efficient data dissemination through optimized replica distribution and achieves superior data collection efficiency under varying node densities and failure scenarios as compared to commonly used uncontrolled sink mobility approaches namely random walk and self-avoiding random walk. By addressing key IoT data management challenges, DRACO offers a scalable and resilient solution well-suited for emerging use cases including industrial IoT device monitoring, smart city environmental sensing, agricultural IoT data collection, and disaster response networks, where maintaining data availability under device failures or intermittent connectivity is critical. Waleed Bin Qaim, Öznur Özkasap, Rabia Qadar, Moncef Gabbouj |
IEEE Internet Things J. | 4 |
| 2026 | GraphCETF: Cost-effective training-free acceleration for evolutionary graph neural architecture search
Bernard-Marie Onzo, Yu Xue 0003, Ferrante Neri, Moncef Gabbouj, Khursheed Aurangzeb |
Knowl. Based Syst. | 4 |
| 2026 | JNTD: Toward Just Noticeable Frame Rate-Based Temporal Difference for Perceptual Video CodingabstractJust Noticeable Difference (JND) refers to the maximum level of distortion in an image or video sequence that remains imperceptible to the Human Visual System (HVS). Current JND-based studies predominantly rely on existing datasets, developing models predicting JND levels in terms of Quantization Parameter (QP) or Quality Factor (QF). However, these solutions primarily focus on spatial-based Perceptual Video Coding (PVC) and neglect temporal-based optimization, which highly affects the video bitrate. This paper addresses this limitation by introducing Just Noticeable frame rate-based Temporal Difference (JNTD) to determine the optimal Frame Rate (FR) based on human perception. A novel dataset comprising 50 high frame rate video sequences is collected through subjective assessments. Subsequently, an ensemble method is proposed to predict the JNTD, by leveraging deep and hand-crafted features, for robust prediction. Experimental evaluations include the integration of the proposed method into several codecs (H.264, H.265, H.266, and a new learned codec), showcasing its ability to reduce bitrate without compromising visual quality. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | DVLTA-VQA: Decoupled Vision-Language Modeling With Text-Guided Adaptation for Blind Video Quality AssessmentabstractInspired by the dual-stream (dorsal and ventral streams) theory of the human visual system (HVS), recent Video Quality Assessment (VQA) methods have integrated Contrastive Language-Image Pretraining (CLIP) to enhance semantic understanding. However, as CLIP is originally designed for images, it lacks the ability to adequately capture the temporal dynamics and motion perception (dorsal stream) inherent in videos. To address this limitation, we propose DVLTA-VQA (Decoupled Vision-Language Modeling with Text-Guided Adaptation), which decouples CLIP’s visual and textual components to better align with the NR-VQA pipeline. Specifically, we introduce a Video-Based Temporal CLIP module and a Temporal Context Module to explicitly model motion dynamics, effectively enhancing the dorsal stream representation. Complementing this, a Basic Visual Feature Extraction Module is employed to strengthen spatial detail analysis in the ventral stream. Furthermore, we propose a text-guided adaptive fusion strategy that leverages textual semantics to dynamically weight visual features, facilitating effective spatiotemporal integration. Extensive experiments on multiple public datasets demonstrate that the proposed method achieves state-of-the-art performance, significantly improving prediction accuracy and generalization capability. Li Yu 0004, Situo Wang, Wei Zhou 0021, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Perception-Inspired Network for Stereo Image Quality AssessmentabstractExisting stereo image quality assessment (SIQA) methods generally have limitations in binocular fusion and fine-grained perception modeling. To address these issues, we propose a Perception-Inspired Network for SIQA that simulates binocular difference-guided fusion, high-frequency sensitivity, and hierarchical perception mechanisms of the human visual system (HVS). First, a difference-guided binocular fusion (DGBF) module is designed to mimic the binocular difference sensitivity mechanism, which exploits difference information at both the feature-level and image-level to optimize binocular fusion. Furthermore, the image distortion primarily affects the high-frequency components, which are critical for perceptual quality. To reflect this, we propose a high-frequency enhancement module (HFEM) to simulate the human eye's sensitivity to edge and texture distortions. Finally, to better achieve fine-grained perception modeling, we propose a hierarchical quality regression strategy that simulates the human perceptual process, from perceiving local details to forming a global quality judgment, thereby achieving a quality prediction more aligned with human subjective evaluation. Experimental results demonstrate that the proposed method outperforms mainstream approaches, achieving a PLCC of 0.9734 on the LIVE I database, and a PLCC of 0.9632 on the LIVE II database. Yongli Chang, Guanghui Yue 0001, Li Yu 0004, Yakun Ju, Hadi Amirpour, Moncef Gabbouj, Wei Zhou 0021 |
IEEE Trans. Image Process. | 7 |
| 2026 | Uncertainty-Guided Spatiotemporal Consistency Fusion Network for Infrared-Visible Video Fusion Under Extremely Low-Light ConditionsabstractInfrared-visible video fusion under extremely low-light conditions is critically important yet remains underexplored, largely due to the scarcity of high-quality datasets and challenges posed by spatiotemporal uncertainty and modality bias. To address the dataset shortage, we built a dataset of 4,739 infrared and visible registration video pairs captured under extremely low-light conditions, spanning 5 scene types and 17 subcategories. Further, we proposed an Uncertainty-guided Spatiotemporal Consistency Fusion Network, termed USCFNet, for the infrared-visible video fusion. At each layer of the encoder, an Entropy-Gated SpatioTemporal Attention (EGSTA) module is introduced to capture temporal instability and spatial reliability variations through entropy-aware attention modulation, thereby enhancing feature spatiotemporal consistency. The refined infrared and visible features are then fused via a Difference-Guided Fusion (DGF) module, which adaptively exploits their content and edge differences to improve structural integrity and detail clarity. By progressively connecting DGF modules from shallow to deep layers, the network achieves the synergistic fusion of shallow textures and deep semantics. Subsequently, the output of the last DGF module is fused with the modality features of the last layer through a hierarchical mixture-of-experts fusion module. This module enables the balanced integration of modality information while preserving fine local details. Finally, the fusion feature is fed into the decoder to produce the final fused video. Extensive experiments on our dataset and two public datasets show that USCFNet outperforms competing methods, achieving lower distortion and stronger spatiotemporal consistency. The source code and dataset are available at https://github.com/Zhaocheng1/ELVID. Cheng Zhao 0003, Tianyun Song, Zhiliang Wu, Tianfu Wang 0001, Moncef Gabbouj, Guanghui Yue 0001, Bai Ying Lei, Wei Zhou 0021 |
IEEE Trans. Image Process. | 5 |
| 2026 | Optimizing the Output of Long Short-Term Memory Cell for High-Frequency Forecasting in Financial MarketsabstractHigh-frequency trading (HFT) requires fast data processing without information lags for precise stock price forecasting. This high-paced stock price forecasting is usually based on vectors that need to be treated as sequential and time-independent signals due to the time irregularities that are inherent in HFT. A well-documented and tested method that considers these time irregularities is a type of recurrent neural network (NN), named long short-term memory (LSTM) NN. This type of NN is formed based on cells that perform sequential and stale calculations via gates and states without knowing whether their order, within the cell, is optimal. In this article, we propose a revised and real-time adjusted LSTM cell that selects the best gate or state as its final output. Our cell is running under a shallow topology, has a minimal look-back period, and is trained online. This revised cell achieves lower forecasting error compared to other recurrent NNs (RNNs) for online HFT forecasting tasks such as the limit order book (LOB) mid-price (MP) prediction as it has been tested on two high-liquid U.S. and two less-liquid Nordic stocks. Adamantios Ntakaris, Moncef Gabbouj, Juho Kanniainen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2026 | A Pairwise Comparison Relation-Assisted Multiobjective Evolutionary Neural Architecture Search Method With Multipopulation MechanismabstractNeural architecture search (NAS) has emerged as a powerful paradigm that enables researchers to automatically explore vast search spaces and discover efficient neural networks. However, NAS suffers from a critical bottleneck, i.e., the evaluation of numerous architectures during the search process demands substantial computing resources and time. In order to improve the efficiency of NAS, a series of methods have been proposed to reduce the evaluation time of neural architectures. However, they are not efficient enough and still only focus on the accuracy of architectures. Beyond classification accuracy, real-world applications increasingly demand more efficient and compact network architectures that balance multiple performance criteria. To address these challenges, we propose the SMEMNAS, a pairwise comparison relation-assisted multiobjective evolutionary algorithm (EA) based on a multipopulation (MP) mechanism. In the SMEMNAS, a surrogate model is constructed based on pairwise comparison relations to predict the accuracy ranking of architectures, rather than the absolute accuracy. Moreover, two populations cooperate with each other in the search process, i.e., a main population that guides the evolutionary process and a vice population that enhances search diversity. Our method aims to discover high-performance models that simultaneously optimize multiple objectives. We conduct comprehensive experiments on CIFAR-10, CIFAR-100, and ImageNet datasets to validate the effectiveness of our approach. With only a single GPU searching for 0.17 days, competitive architectures can be found by SMEMNAS, which achieves 78.91% accuracy with the MAdds of 570 M on the ImageNet. This work makes a significant advancement in the field of NAS. Yu Xue 0003, Pengcheng Jiang, Chenchen Zhu, MengChu Zhou, Mohamed Wahib, Moncef Gabbouj |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2025 | To Vaccinate or Not to Vaccinate? Analyzing $\mathbb {X}$ Power over the Pandemic
Tanveer Khan, Fahad Sohrab, Antonis Michalas, Moncef Gabbouj |
AINA (7) | 4 |
| 2025 | Hierarchical Transformer for Panoramic Image Inpainting with Comprehensive Attention Module
Li Yu 0004, Yanjun Gao, Yihang Yin, Farhad Pakdaman, Moncef Gabbouj |
ICIG (2) | 5 |
| 2025 | Evaluating the Emerging MPEG Video Coding for Machines in Semantic SegmentationabstractEmerging MPEG Video Coding for Machines (MPEG VCM) standardization activities address the growing demand for machine-to-machine visual applications, including video surveillance, autonomous driving, etc. This paper proposes an evaluation methodology tailored to MPEG VCM, with an emphasis on semantic segmentation tasks using the Pandaset dataset. This is a challenging target as standardization works must follow several limitations, such as dataset's licensing, fixed tools and software, and compliance with existing common test conditions (CTC) for the standard's development. The proposed evaluation methodology includes a step to align Pandaset and COCO labels. Two task networks, Detectron2 and Mask2Former, are used to evaluate the Rate-Performance behavior for semantic segmentation under various coding configurations. The performance of MPEG VCM is benchmarked against traditional codecs (VVC and HEVC), with a detailed analysis of MPEG VCM's coding tools. The extensive evaluations reveal interesting observations. (1) Although VCM achieved reasonably good segmentation performance, some of its developed tools, such as temporal resampling and region-of-interest coding, were not well suited for segmentation task. (2) The Hybrid NNVVC Inner Codec outperformed the VVC Inner Codec. (3) VCM's performance varies significantly for segmented classes. (4) Despite significant differences in human vision performance, VVC and HEVC exhibit relatively similar performance in machine vision. The main contributions of this work are to (1) enable evaluation of MPEG VCM in a real-world semantic segmentation use case, which is one of VCM's targeted tasks, and (2) to provide a detailed assessment of VCM's performance in semantic segmentation. Khoa Dang Pham, Farhad Pakdaman, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Nam Le 0003, Jukka I. Ahonen, Moncef Gabbouj |
ISM | 7 |
| 2025 | Randomized PCA forest for approximate k-nearest neighbor searchabstractk-Nearest Neighbors (kNN) search is the problem of finding k points which are the closest to a given query point . It is used widely in a wide range of tasks and is among the most important tools in applied machine learning . Traditional algorithms for kNN search require computing distances between a query point and all other points in the dataset, and therefore is very slow and inefficient for large data. In this paper, we propose an approximate algorithm for kNN search to find the nearest neighbors fast and efficiently. We employ a tree-based structure which offers robustness and scalability. We propose to use Principal Component Analysis (PCA) to find the best splitting direction to fit the data on the trees. Seeking solutions with low computational complexity , (1) we use a randomized Singular Value Decomposition solver, which reduces PCA complexity from being associated with the number of features to being associated with the number of required principal values; (2) we reuse PCA calculations in multiple nodes to save computation while maintaining accuracy; (3) we ensemble these trees for improved performance, and (4) finally, we propose several variants of the proposed method which target a higher accuracy or a higher efficiency. Extensive experimental results show that proposed solutions outperform existing methods in terms of accuracy, while maintaining competitive complexity. The fast implementation variant of the proposed method outperforms existing techniques in terms of complexity and shows competitive accuracy in performing k-nearest neighbors’ search. Muhammad Rajabinasab, Farhad Pakdaman, Arthur Zimek, Moncef Gabbouj |
Expert Syst. Appl. | 4 |
| 2025 | High-Frequency Enhanced Hybrid Neural Representation for video compression
Li Yu 0004, Jimin Xiao, Moncef Gabbouj |
Expert Syst. Appl. | 4 |
| 2025 | YOLO-DKR: Differentiable architecture search based on kernel reusing for object detection
Yu Xue 0003, Chenhang Yao, Mohamed Wahib, Moncef Gabbouj |
Inf. Sci. | 4 |
| 2025 | Dropout Concrete Autoencoder for Band Selection on Hyperspectral Image ScenesabstractDeep learning-based informative band selection methods on hyperspectral images (HSI) have recently gained intense attention to eliminate spectral correlation and redundancies. However, existing deep learning-based methods either need additional post-processing strategies to select the descriptive bands or optimize the model indirectly due to the parameterization inability of discrete variables for the selection procedure. To overcome these limitations, this work proposes a novel end-to-end network for informative band selection. The proposed network, named Dropout CAE, is inspired by advances in the concrete autoencoder (CAE) and dropout feature ranking (Dropout FR) strategy. Unlike traditional deep learning-based methods; the Dropout CAE is trained directly given the required band subset, eliminating the need for further post-processing. Experimental results in four HSI scenes show that the Dropout CAE achieves substantial and effective performance levels that outperform competing methods. The code is available at https: //github.com/LeiXuAI/Hyperspectral. Lei Xu 0036, Mete Ahishali, Moncef Gabbouj |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | BRSR-OpGAN: Blind radar signal restoration using operational generative adversarial networkabstractMany studies on radar signal restoration in the literature focus on isolated restoration problems, such as denoising over a certain type of noise, while ignoring other types of artifacts. Additionally, these approaches usually assume a noisy environment with a limited set of fixed signal-to-noise ratio (SNR) levels. However, real-world radar signals are often corrupted by a blend of artifacts, including but not limited to unwanted echo, sensor noise, intentional jamming, and interference, each of which can vary in type, severity, and duration. This study introduces Blind Radar Signal Restoration using an Operational Generative Adversarial Network (BRSR-OpGAN), which uses a dual domain loss in the temporal and spectral domains. This approach is designed to improve the quality of radar signals, regardless of the diversity and intensity of the corruption. The BRSR-OpGAN utilizes 1D Operational GANs, which use a generative neuron model specifically optimized for blind restoration of corrupted radar signals. This approach leverages GANs' flexibility to adapt dynamically to a wide range of artifact characteristics. The proposed approach has been extensively evaluated using a well-established baseline and a newly curated extended dataset called the Blind Radar Signal Restoration (BRSR) dataset. This dataset was designed to simulate real-world conditions and includes a variety of artifacts, each varying in severity. The evaluation shows an average SNR improvement over 15.1 dB and 14.3 dB for the baseline and BRSR datasets, respectively. Finally, the proposed approach can be applied in real-time, even on resource-constrained platforms. This pilot study demonstrates the effectiveness of blind radar restoration in time-domain for real-world radar signals, achieving exceptional performance across various SNR values and artifact types. The BRSR-OpGAN method exhibits robust and computationally efficient restoration of real-world radar signals, significantly outperforming existing methods. Muhammad Uzair Zahid, Serkan Kiranyaz, Alper Yildirim, Moncef Gabbouj |
Neural Networks | 4 |
| 2025 | Deep-BrownConrady: Prediction of Camera Calibration and Distortion Parameters Using Deep Learning and Synthetic DataabstractThis research addresses the challenge of camera calibration and distortion parameter prediction from a single image using deep learning models. The main contributions of this work are: (1) demonstrating that a deep learning model, trained on a mix of real and synthetic images, can accurately predict camera and lens parameters from a single image, and (2) developing a comprehensive synthetic dataset using the AILiveSim simulation platform. This dataset includes variations in focal length and lens distortion parameters, providing a robust foundation for model training and testing. The training process predominantly relied on these synthetic images, complemented by a small subset of real images, to explore how well models trained on synthetic data can perform calibration tasks on real-world images. Traditional calibration methods require multiple images of a calibration object from various orientations, which is often not feasible due to the lack of such images in publicly available datasets. A deep learning network based on the ResNet architecture was trained on this synthetic dataset to predict camera calibration parameters following the Brown-Conrady lens model. The ResNet architecture, adapted for regression tasks, is capable of predicting continuous values essential for accurate camera calibration in applications such as autonomous driving, robotics, and augmented reality. Faiz Muhammad Chaudhry, Jarno Ralli, Jérôme Leudet, Fahad Sohrab, Farhad Pakdaman, Pierre Corbani, Moncef Gabbouj |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | Integrating Recurrent-KAN With SAM Adapter for Blind Hyperspectral UnmixingabstractDue to the limitation of sensors, hyperspectral images contain a large number of mixed pixels. Hyperspectral unmixing techniques decompose these mixed pixels into distinct endmembers and their corresponding abundance values. Traditional methods initialize weights in the decoder and utilize outputs/weights as abundance maps and endmembers——an approach heavily dependent on initial weights that significantly limits performance. This paper proposes a blind hyperspectral unmixing method integrating Recurrent Kolmogorov-Arnold Networks (KAN) with Segment Anything Model (SAM) adapter. The method operates through three sequential stages: feature encoding, endmember extraction, and abundance estimation. Specifically for feature encoding, a HU-SAM adapter is proposed to capture global-local spatial features. For endmember extraction, an iteratively learned Recurrent-KAN module reconstructs endmembers while stabilizing model learning. For abundance estimation, an updated Swin Transformer module is utilized to maintain a lower parameter count. Extensive experiments on real and synthetic datasets demonstrate superior effectiveness of the proposed method over the eight state-of-the-art methods. Yihao Fu, Shenglin Peng, Jun Wang 0078, Jinye Peng 0001, Moncef Gabbouj |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Crucial-Diff: A Unified Diffusion Model for Crucial Image and Annotation Synthesis in Data-Scarce ScenariosabstractThe scarcity of data in various scenarios, such as medical, industry and autonomous driving, leads to model overfitting and dataset imbalance, thus hindering effective detection and segmentation performance. Existing studies employ the generative models to synthesize more training samples to mitigate data scarcity. However, these synthetic samples are repetitive or simplistic and fail to provide "crucial information" that targets the downstream model's weaknesses. Additionally, these methods typically require separate training for different objects, leading to computational inefficiencies. To address these issues, we propose Crucial-Diff, a domain-agnostic framework designed to synthesize crucial samples. Our method integrates two key modules. The Scene Agnostic Feature Extractor (SAFE) utilizes a unified feature extractor to capture target information. The Weakness Aware Sample Miner (WASM) generates hard-to-detect samples using feedback from the detection results of downstream model, which is then fused with the output of SAFE module. Together, our Crucial-Diff framework generates diverse, high-quality training data, achieving a pixel-level AP of 83.63% and an F1-MAX of 78.12% on MVTec. On polyp dataset, Crucial-Diff reaches an mIoU of 81.64% and an mDice of 87.69%. Code is publicly available at https://github.com/JJessicaYao/Crucial-diff. Siyue Yao, Mingjie Sun, Eng Gee Lim, Ran Yi 0002, Baojiang Zhong, Moncef Gabbouj |
IEEE Trans. Image Process. | 6 |
| 2025 | Introduction to the Special Issue on AI Empowered Edge Computing for Multimedia ApplicationsabstractNo abstract available. Moncef Gabbouj, Jin Li 0002, Haibo Hu 0001, Yang Xiang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | A Survey on Securing Image-Centric Edge IntelligenceabstractFacing enormous data generated at the network edge, edge intelligence (EI) emerges as the fusion of edge computing and AI, revolutionizing edge data processing and intelligent decision-making. Nonetheless, this emergent mode presents a complex array of security challenges, particularly prominent in image-centric applications due to the sheer volume of visual data and its direct connection to user privacy. These challenges include safeguarding model/image privacy and ensuring model integrity against various security threats, such as model poisoning. Essentially, those threats originate from data attacks, suggesting data protection as a promising solution. Although data protection measures are well-established in other domains, image-centric EI necessitates focused research. This survey examines the security issues inherent to image-centric EI and outlines the protection efforts, providing a comprehensive overview of the landscape. We begin by introducing EI, detailing its operational mechanics and associated security issues. We then explore the technologies facilitating security enhancement (e.g., differential privacy) and EI (e.g., compact networks and distributed learning frameworks). Next, we categorize security strategies by their application in data preparation, training, and inference, with a focus on image-based contexts. Despite these efforts on security, our investigation identifies research gaps. We also outline promising research directions to bridge these gaps, bolstering security frameworks in image-centric EI applications. Haibo Hu 0001, Moncef Gabbouj, Qingqing Ye 0001, Yang Xiang 0001, Jin Li 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Trustworthiness of $\mathbb {X}$ Users: A One-Class Classification Approach
Tanveer Khan, Fahad Sohrab, Antonis Michalas, Moncef Gabbouj |
AINA (2) | 4 |
| 2024 | Panoramic Image Inpainting with Gated Convolution and Contextual Reconstruction LossabstractDeep learning-based methods have demonstrated encouraging results in tackling the task of panoramic image inpainting. However, it is challenging for existing methods to distinguish valid pixels from invalid pixels and find suitable references for corrupted areas, thus leading to artifacts in the inpainted results. In response to these challenges, we propose a panoramic image inpainting framework that consists of a Face Generator, a Cube Generator, a side branch, and two discriminators. We use the Cubemap Projection (CMP) format as network input. The generator employs gated convolutions to distinguish valid pixels from invalid ones, while a side branch is designed utilizing contextual reconstruction (CR) loss to guide the generators to find the most suitable reference patch for inpainting the missing region. The proposed method is compared with state-of-the-art (SOTA) methods on SUN360 Street View dataset in terms of PSNR and SSIM. Experimental results and ablation study demonstrate that the proposed method outperforms SOTA both quantitatively and qualitatively. Li Yu 0004, Yanjun Gao, Farhad Pakdaman, Moncef Gabbouj |
ICASSP | 4 |
| 2024 | Uimt: A Framework for Improving Unimodal Inference via Multimodal TrainingabstractThe field of multimodal learning is developing rapidly, with emergence of many novel models and applications. Still, works proposing unimodal and multimodal models are generally disjoint, and either focus on fully-unimodal or fully-multimodal scenarios. Nevertheless, oftentimes in real-world applications data of multiple modalities are available during training while only one of them can be utilized during inference due to associated computational costs, or complexity of utilizing additional sensors. In this work, we develop a framework for improving inference of arbitrary unimodal models with multimodal training, without incurring any additional computational cost at inference time, but benefiting from the advantages of multimodal training. We show that our framework is applicable to different architecture types: transformers, 3D CNNs, and 2D+1D CNNs. To showcase this generality we evaluate our approach on tasks of dynamic hand gesture recognition based on RGB and Depth, audiovisual emotion recognition based, and audio-video-text based sentiment analysis. Our approach consistently outperforms the conventionally trained unimodal counterparts. We additionally investigate how within our framework training of multimodal models can benefit from unimodal, modality-specific learning signals. Utilizing the same variety of architectures as mentioned above, we show how models trained with additional supervision from each isolated modality outperform a multimodal-only counterpart. Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj |
ICIP | 3 |
| 2024 | Pixel-Wise Color Constancy Via Smoothness Techniques In Multi-Illuminant ScenesabstractMost scenes are illuminated by several light sources, where the traditional assumption of uniform illumination is invalid. This issue is ignored in most color constancy methods, primarily due to the complex spatial impact of multiple light sources on the image. Moreover, most existing multi-illuminant methods fail to preserve the smooth change of illumination, which stems from spatial dependencies in natural images. Motivated by this, we propose a novel multi-illuminant color constancy method, by learning pixel-wise illumination maps caused by multiple light sources. The proposed method enforces smoothness within neighboring pixels, by regularizing the training with the total variation loss. Moreover, a bilateral filter is provisioned further to enhance the natural appearance of the estimated images, while preserving the edges. Additionally, we propose a label-smoothing technique that enables the model to generalize well despite the uncertainties in ground truth. Quantitative and qualitative experiments demonstrate that the proposed method outperforms the state-of-the-art. Umut Cem Entok, Firas Laakom, Farhad Pakdaman, Moncef Gabbouj |
ICIP | 4 |
| 2024 | Perceptual Learned Image Compression via End-to-End JND-Based OptimizationabstractEmerging Learned image Compression (LC) achieves significant improvements in coding efficiency by end-to-end training of neural networks for compression. An important benefit of this approach over traditional codecs is that any optimization criteria can be directly applied to the encoder-decoder networks during training. Perceptual optimization of LC to comply with the Human Visual System (HVS) is among such criteria, which has not been fully explored yet. This paper addresses this gap by proposing a novel framework to integrate Just Noticeable Distortion (JND) principles into LC. Leveraging existing JND datasets, three perceptual optimization methods are proposed to integrate JND into the LC training process: (1) Pixel-Wise JND Loss (PWL) prioritizes pixel-by-pixel fidelity in reproducing JND characteristics, (2) Image-Wise JND Loss (IWL) emphasizes on overall imperceptible degradation levels, and (3) Feature-Wise JND Loss (FWL) aligns the reconstructed image features with perceptually significant features. Experimental evaluations demonstrate the effectiveness of JND integration, highlighting improvements in rate-distortion performance and visual quality, compared to baseline methods. The proposed methods add no extra complexity after training. Farhad Pakdaman, Sanaz Nami, Moncef Gabbouj |
ICIP | 3 |
| 2024 | Joint End-to-End Image Compression and Denoising: Leveraging Contrastive Learning and Multi-Scale Self-OnnsabstractNoisy images are a challenge to image compression algorithms due to the inherent difficulty of compressing noise. As noise cannot easily be discerned from image details, such as high-frequency signals, its presence leads to extra bits needed for compression Since the emerging learned image compression paradigm enables end-to-end optimization of codecs, recent efforts were made to integrate denoising into the compression model, relying on clean image features to guide denoising. However, these methods exhibit suboptimal performance under high noise levels, lacking the capability to generalize across diverse noise types. In this paper, we propose a novel method integrating a multi-scale denoiser comprising of Self Organizing Operational Neural Networks, for joint image compression and denoising. We employ contrastive learning to boost the network ability to differentiate noise from high frequency signal components, by emphasizing the correlation between noisy and clean counterparts. Experimental results demonstrate the effectiveness of the proposed method both in rate-distortion performance, and codec speed, outperforming the current state-of-the-art. Li Yu 0004, Farhad Pakdaman, Moncef Gabbouj |
ICIP | 4 |
| 2024 | Intrinsic Image Decomposition Based on Quantized Prior CodebookabstractIntrinsic image decomposition is a low-level image processing task that extracts the reflectance and lighting components from an image. This process can improve the illumination robustness of perception tasks, such as object detection, recognition, and image understanding. Recently, deep image generation frameworks have been used to generate intrinsic images. However, the encoder and decoder lack prior knowledge constraints. This paper presents a quantized codebook for embedding intrinsic features that guide the extraction of intrinsic images. To enhance reconstruction accuracy, we propose a purification method to eliminate irrelevant elements from the codebook. Additionally, we propose self-attention and cross-attention modules to integrate the intrinsic features of the codebook into the input image features for reconstruction. The effectiveness of the algorithm is demonstrated through experiments conducted on several popular datasets. Fangzheng Yuan, Xiaoyue Jiang, Xiaoyi Feng, Moncef Gabbouj |
ICIP | 4 |
| 2024 | Refining Myocardial Infarction Detection: A Novel Multi-Modal Composite Kernel Strategy in One-Class ClassificationabstractEarly detection of myocardial infarction (MI), a critical condition arising from coronary artery disease (CAD), is vital to prevent further myocardial damage. This study introduces a novel method for early MI detection using a one-class classification (OCC) algorithm in echocardiography. Our study overcomes the challenge of limited echocardiography data availability by adopting a novel approach based on Multi-modal Subspace Support Vector Data Description. The proposed technique involves a specialized MI detection framework employing multi-view echocardiography incorporating a composite kernel in the non-linear projection trick, fusing Gaussian and Laplacian sigmoid functions. Additionally, we enhance the update strategy of the projection matrices by adapting maximization for both or one of the modalities in the optimization process. Our method boosts MI detection capability by efficiently transforming features extracted from echocardiography data into an optimized lower-dimensional subspace. The OCC model trained specifically on target class instances from the comprehensive HMC-QU dataset that includes multiple echocardiography views indicates a marked improvement in MI detection accuracy. Our findings reveal that our proposed multi-view approach achieves a geometric mean of 71.24%, signifying a substantial advancement in echocardiography-based MI diagnosis and offering more precise and efficient diagnostic tools. Muhammad Uzair Zahid, Aysen Degerli, Fahad Sohrab, Serkan Kiranyaz, Tahir Hamid, Rashid Mazhar, Moncef Gabbouj |
ICIP | 7 |
| 2024 | Water Region Segmentation in SAR Images Based on Compact Operational UnetsabstractIn this work, we propose a novel supervised segmentation approach based on compact Operational U-Nets (Op-UNets) for detection of water regions in high-resolution synthetic aperture radar (SAR) images. The proposed approach utilizes the Segment Anything Model (SAM) for supervised training of separate compact Op-UNet model to automatically generate higher accuracy binary segmentation masks from the ground truth optical imagery. To the best of our knowledge, this is the first study that applies jointly the SAM model and compact Op-UNet architecture which utilizes the self-organized operational neural network (self-ONN) layers to improve SAR image segmentation performance with limited labeled data. The preliminary results from the performed experiments using the high-resolution SAR imagery over the coastline along Daytona Beach and Seminole, FL (from Capella Space) are presented to evaluate the performance of the proposed model with significantly reduced computational complexity. Turker Ince, Steven Beninati, Ozer Can Devecioglu, Stephen J. Frasier, Moncef Gabbouj |
IGARSS | 5 |
| 2024 | MAESR360: Masked autoencoder-based 360-degree video streaming via multi-scale feature fusionabstract360-degree video streaming is becoming increasingly popular for its immersive experience. Traditional adaptive tile-based streaming methods allocate the bitrates according to view-port prediction, which effectively reduces required transmission bandwidth, but it will cause serious quality degradation when the viewport prediction is inaccurate. Thus, some researchers propose visual reconstruction and enhancement-based 360-degree video streaming framework, which can reconstructs the whole frame at very low bitrates. However, existing frameworks are built upon image-based visual reconstruction methods, which do not fully consider the characteristics of videos. In this paper, we propose a masked autoencoder-based, multi-scale optimized framework for 360-degree video streaming (MAESR360), which fully considers the temporal relevance of the video. We utilize spatio-temporal downsampling and high-ratio tube masking strategies to effectively reduce the amount of transmitted data. Additionally, we design a lightweight visual reconstruction model based on multi-scale feature fusion to recover the visual quality of video frames. The effectiveness of our proposed method is demonstrated through extensive experiments. Li Yu 0004, Zhiyu Pang, Moncef Gabbouj |
VCIP | 3 |
| 2024 | Operational Support Estimator NetworksabstractIn this work, we propose a novel approach called Operational Support Estimator Networks (OSENs) for the support estimation task. Support Estimation (SE) is defined as finding the locations of non-zero elements in sparse signals. By its very nature, the mapping between the measurement and sparse signal is a non-linear operation. Traditional support estimators rely on computationally expensive iterative signal recovery techniques to achieve such non-linearity. Contrary to the convolutional layers, the proposed OSEN approach consists of operational layers that can learn such complex non-linearities without the need for deep networks. In this way, the performance of non-iterative support estimation is greatly improved. Moreover, the operational layers comprise so-called generative super neurons with non-local kernels. The kernel location for each neuron/feature map is optimized jointly for the SE task during training. We evaluate the OSENs in three different applications: i. support estimation from Compressive Sensing (CS) measurements, ii. representation-based classification, and iii. learning-aided CS reconstruction where the output of OSENs is used as prior knowledge to the CS algorithm for enhanced reconstruction. Experimental results show that the proposed approach achieves computational efficiency and outperforms competing methods, especially at low measurement rates by significant margins. Mete Ahishali, Mehmet Yamac, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | R2C-GAN: Restore-to-Classify Generative Adversarial Networks for blind X-ray restoration and COVID-19 classificationabstractRestoration of poor-quality medical images with a blended set of artifacts plays a vital role in a reliable diagnosis. As a pioneer study in blind X-ray restoration, we propose a joint model for generic image restoration and classification: Restore-to-Classify Generative Adversarial Networks (R2C-GANs). This is the first generic restoration approach forming an Image-to-Image translation task from poor-quality having noisy, blurry, or over/under-exposed images to high-quality image domain where forward and inverse transformations are learned using unpaired training samples. Simultaneously, the joint classification preserves the diagnostic-related label during restoration. Each R2C-GAN is equipped with operational layers/neurons in a compact architecture. The proposed joint model successfully restores images while achieving state-of-the-art Coronavirus Disease 2019 (COVID-19) classification with above 90% in F1-Score. In qualitative analysis, the restoration performance is confirmed by medical doctors where 68% of the restored images are selected against the original images. We share the software implementation at https://github.com/meteahishali/R2C-GAN. Mete Ahishali, Aysen Degerli, Serkan Kiranyaz, Tahir Hamid, Rashid Mazhar, Moncef Gabbouj |
Pattern Recognit. | 6 |
| 2024 | Reducing redundancy in the bottleneck representation of autoencodersabstractAutoencoders (AEs) are a type of unsupervised neural networks, which can be used to solve various tasks, e.g., dimensionality reduction, image compression, and image denoising. An AE has two goals: (i) compress the original input to a low-dimensional space at the bottleneck of the network topology using an encoder, (ii) reconstruct the input from the representation at the bottleneck using a decoder. Both encoder and decoder are optimized jointly by minimizing a distortion-based loss which implicitly forces the model to keep only the information in input data required to reconstruct them and to reduce redundancies. In this paper, we propose a scheme to explicitly penalize feature redundancies in the bottleneck representation. To this end, we propose an additional loss term, based on the pairwise covariances of the network units, which complements the data reconstruction loss forcing the encoder to learn a more diverse and richer representation of the input. We tested our approach across different tasks, namely dimensionality reduction, image compression, and image denoising. Experimental results show that the proposed loss leads consistently to superior performance compared to using the standard AE loss. Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. Lett. | 4 |
| 2024 | Channel-Wise Feature Decorrelation for Enhanced Learned Image CompressionabstractThe emerging Learned Compression (LC) replaces the traditional codec modules with Deep Neural Networks (DNN), which are trained end-to-end for rate-distortion performance. This approach is considered as the future of image/video compression, and major efforts have been dedicated to improving its compression efficiency. However, most proposed works target compression efficiency by employing more complex DNNS, which contributes to higher computational complexity. Alternatively, this paper proposes to improve compression by fully exploiting the existing DNN capacity. To do so, the latent features are guided to learn a richer and more diverse set of features, which corresponds to better reconstruction. A channel-wise feature decorrelation loss is designed and is integrated into the LC optimization. Three strategies are proposed and evaluated, which optimize (1) the transformation network, (2) the context model, and (3) both networks. Experimental results on two established LC methods show that the proposed method improves the compression with a BD-Rate of up to 8.06%, with no added complexity. The proposed solution can be applied as a plug-and-play solution to optimize any similar LC method. Farhad Pakdaman, Moncef Gabbouj |
IEEE Signal Process. Lett. | 2 |
| 2024 | Multi-Swin Transformer Based Spatio-Temporal Information Exploration for Compressed Video Quality EnhancementabstractSpatio-temporal information plays an important role in compressed video quality enhancement. Most advanced studies use deformable convolution or Swin transformer to explore spatio-temporal information. However, deformable convolution based methods may incur inaccurate motion compensation due to the compression artifacts and limited receptive fields. The Swin transformer based approaches are unable to fully explore the spatio-temporal information, limited by its rigid window-based mechanism. To solve the above problems, we propose a novel multi-Swin transformer-based network for compressed video quality enhancement to better explore spatio-temporal information. The whole workflow consists of the Local Alignment (LA) Module, the Global Refinement Fusion (GRF) Module, and the Quality Enhancement (QE) Module. The LA module roughly perceives the local motion through the deformable fusion. Subsequently, the GRF module employs the proposed multi-Swin transformer to enhance the spatio-temporal perception. Finally, the QE module effectively restores the texture details across various scales. Extensive experimental results prove the effectiveness of the proposed method. Li Yu 0004, Shiyu Wu, Moncef Gabbouj |
IEEE Signal Process. Lett. | 3 |
| 2024 | Lightweight Multitask Learning for Robust JND Prediction Using Latent Space and Reconstructed FramesabstractThe Just Noticeable Difference (JND) refers to the smallest distortion in an image or video that can be perceived by Human Visual System (HVS), and is widely used in optimizing image/video compression. However, accurate JND modeling is very challenging due to its content dependence, and the complex nature of the HVS. Recent solutions train deep learning based JND prediction models, mainly based on a Quantization Parameter (QP) value, representing a single JND level, and train separate models to predict each JND level. We point out that a single QP-distance is insufficient to properly train a network with millions of parameters, for a complex content-dependent task. Inspired by recent advances in learned compression and multitask learning, we propose to address this problem by (1) learning to reconstruct the JND-quality frames, jointly with the QP prediction, and (2) jointly learning several JND levels to augment the learning performance. We propose a novel solution where first, an effective feature backbone is trained by learning to reconstruct JND-quality frames from the raw frames. Second, JND prediction models are trained based on features extracted from latent space (i.e., compressed domain), or reconstructed JND-quality frames. Third, a multi-JND model is designed, which jointly learns three JND levels, further reducing the prediction error. Extensive experimental results demonstrate that our multi-JND method outperforms the state-of-the-art and achieves an average JND1prediction error of only 1.57 in QP, and 0.72 dB in PSNR. Moreover, the multitask learning approach, and compressed domain prediction facilitate light-weight inference by significantly reducing the complexity and the number of parameters. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | WLD-Reg: A Data-Dependent Within-Layer Diversity RegularizerabstractNeural networks are composed of multiple layers arranged in a hierarchical structure jointly trained with a gradient-based optimization, where the errors are back-propagated from the last layer back to the first one. At each optimization step, neurons at a given layer receive feedback from neurons belonging to higher layers of the hierarchy. In this paper, we propose to complement this traditional 'between-layer' feedback with additional 'within-layer' feedback to encourage the diversity of the activations within the same layer. To this end, we measure the pairwise similarity between the outputs of the neurons and use it to model the layer's overall diversity. We present an extensive empirical study confirming that the proposed approach enhances the performance of several state-of-the-art neural network models in multiple tasks. The code is publically available at https://github.com/firasl/AAAI-23-WLD-Reg. Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
AAAI | 4 |
| 2023 | Comprehensive Complexity Assessment of Emerging Learned Image Compression on CPU and GPUabstractLearned Compression (LC) is the emerging technology for compressing image and video content, using deep neural networks. Despite being new, LC methods have already gained a compression efficiency comparable to state-of-the-art image compression, such as HEVC or even VVC. However, the existing solutions often require a huge computational complexity, which discourages their adoption in international standards or products. This paper provides a comprehensive complexity assessment of several notable methods, that shed light on the matter, and guide the future development of this field by presenting key findings. To do so, six existing methods have been evaluated for both encoding and decoding, on CPU and GPU platforms. Various aspects of complexity such as the overall complexity, share of each coding module, number of operations, number of parameters, most demanding GPU kernels, and memory requirements have been measured and compared on Kodak dataset. The reported results (1) quantify the complexity of LC methods, (2) fairly compare different methods, and (3) a major contribution of the work is identifying and quantifying the key factors affecting the complexity. Farhad Pakdaman, Moncef Gabbouj |
ICASSP | 2 |
| 2023 | MTJND: Multi-Task Deep Learning Framework for Improved JND PredictionabstractThe limitation of the Human Visual System (HVS) in perceiving small distortions allows us to lower the bitrate required to achieve a certain visual quality. Predicting and applying the Just Noticeable Distortion (JND), which is a threshold for maximum unperceived level of distortions, is among the popular ways to do so. Recently, machine learning based methods have been able to reduce bitrate even further by improving JND prediction accuracy. However, accurate modeling of JND is very challenging, as it is highly content dependent. Furthermore, existing datasets provide little information to learn the best parameters. To remedy this issue, we propose a multi-task deep learning framework that jointly learns various complementary visual information. We design three separate methods and training strategies that jointly learn: (1) three JND levels, (2) visual attention map and a JND level, and (3) three JND levels and the visual attention map. We show that accumulating information from multiple tasks leads to a more robust prediction of JND. Experimental results confirm the superiority of our framework compared to the state-of-the-art. Sanaz Nami, Farhad Pakdaman, Mahmoud Reza Hashemi, Shervin Shirmohammadi, Moncef Gabbouj |
ICIP | 5 |
| 2023 | NCOD: Near-Optimum Video Compression for Object DetectionabstractWith the emergence of technologies like smart cities, Internet of things (IoT), and 5G, the amount of produced visual data at the edges and remote nodes has exploded. Since for a considerable portion of the captured video the target is a machine learning task, rather than a human audience, transmission of videos in such applications requires efficient video compression tailored for machine vision. However, existing compression solutions are optimized for human vision. This paper presents a methodology to optimize an existing video compression standard, HEVC, for a machine vision task, Object Detection (OD). To this end, (1) a dataset of compressed videos, including several compression-ratios and their corresponding OD performance is collected to enable modeling, (2) A trade-off point (knee-point) between bitrate and OD performance is defined, that finds the point after which no major improvements will be achieved, (3) a set of features were extracted and studied to model this point, via a practical machine learning method. The resulting solution can predict the knee-point with$\boldsymbol{\text{MAE}=1.28}$, resulting in a ∆Recall of only 0.012 and bitrate reduction of 86.56%, compared to OD with very high-quality video. Ardavan Elahi, Ali Falahati, Farhad Pakdaman, Mehdi Modarressi, Moncef Gabbouj |
ISCAS | 5 |
| 2023 | Optimal Tile Size and Streaming Field of View for VR StreamingabstractVirtual reality (VR) video services require a high bitrate, and hence, viewport-adaptive streaming techniques like motion-constrained-tile-set (MCTS) have been found important to reduce streaming-rate and storage demands. The tiling scheme and streaming field of view (FOV) are among the key elements in designing an optimal VR viewport-adaptive streaming solution, in terms of rate-distortion (R-D) performance. The aim of this study is to propose an optimal configuration for the tile grid and streaming FOV, considering different VR viewing situations such as head motion speed, system delay, and head-mounted display FOV. To achieve this, a wide range of tiling schemes and streaming FOVs are examined to study the storage and streaming R-D performance of the MCTS-based technique in both viewport and non-viewport areas using a quality metric called Zonal-cubic PSNR. The findings demonstrate that for VR applications focused on preserving high viewport quality, fine tile grids lead to higher performance. In scenarios featuring small and large HMD FOV, the optimal configuration involves a small and medium streaming FOV, respectively. Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
MMSP | 4 |
| 2023 | Learning Distinct Features Helps, Provably
Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
ECML/PKDD (2) | 4 |
| 2023 | An external attention-based feature ranker for large-scale feature selectionabstractAn important problem in data science, feature selection (FS) consists of finding the optimal subset of features and eliminating irrelevant or redundant features. The FS task on high-dimensional data is challenging for the FS methods currently available in the literature. To overcome this limitation, we propose a novel feature selection method called External Attention-Based Feature Ranker for Large-Scale Feature Selection (EAR-FS) whose function is based on the logic of an attention mechanism and a hybrid metaheuristic. EAR-FS comprises three interdependent modules: (1) in the training module design, a multilayer perceptron network endowed with an attention module is trained to fit the dataset; (2) in feature ranking by attention, the trained attention module is used for attention updating and to rank features according to their importance; 3) in subset generation, a two-stage heuristic approach is applied to determine a small number of features that still guarantee high-accuracy performance. The experimental benchmark comprised 26 datasets of small, large and very large sizes, ranging from 15 to 12,533 features. Experiments performed against the state-of-the-art algorithms of FS show that our algorithm is efficient at selecting a small number of features from large datasets while guaranteeing excellent levels of classification accuracy. For instance, EAR-FS demonstrated its capability to reduce the features of the 11 Tumor dataset by 97% while maintaining a classifier accuracy of over 93%. Yu Xue 0003, Ferrante Neri, Moncef Gabbouj, Yong Zhang 0016 |
Knowl. Based Syst. | 4 |
| 2023 | Representation based regression for object distance estimationabstractIn this study, we propose a novel approach to predict the distances of the detected objects in an observed scene. The proposed approach modifies the recently proposed Convolutional Support Estimator Networks (CSENs). CSENs are designed to compute a direct mapping for the Support Estimation (SE) task in a representation-based classification problem. We further propose and demonstrate that representation-based methods (sparse or collaborative representation) can be used in well-designed regression problems especially over scarce data. To the best of our knowledge, this is the first representation-based method proposed for performing a regression task by utilizing the modified CSENs; and hence, we name this novel approach as Representation-based Regression (RbR). The initial version of CSENs has a proxy mapping stage (i.e., a coarse estimation for the support set) that is required for the input. In this study, we improve the CSEN model by proposing Compressive Learning CSEN (CL-CSEN) that has the ability to jointly optimize the so-called proxy mapping stage along with convolutional layers. The experimental evaluations using the KITTI 3D Object Detection distance estimation dataset show that the proposed method can achieve a significantly improved distance estimation performance over all competing methods. Finally, the software implementations of the methods are publicly shared at https://github.com/meteahishali/CSENDistance. Mete Ahishali, Mehmet Yamac, Serkan Kiranyaz, Moncef Gabbouj |
Neural Networks | 4 |
| 2023 | Joint learning and optimization for Federated Learning in NOMA-based networks
Ilyes Mrad, Ridha Hamila, Aiman Erbad, Moncef Gabbouj |
Pervasive Mob. Comput. | 4 |
| 2023 | Graph-embedded subspace support vector data descriptionabstractIn this paper, we propose a novel subspace learning framework for one-class classification. The proposed framework presents the problem in the form of graph embedding. It includes the previously proposed subspace one-class techniques as its special cases and provides further insight on what these techniques actually optimize. The framework allows to incorporate other meaningful optimization goals via the graph preserving criterion and reveals a spectral solution and a spectral regression-based solution as alternatives to the previously used gradient-based technique. We combine the subspace learning framework iteratively with Support Vector Data Description applied in the subspace to formulate Graph-Embedded Subspace Support Vector Data Description. We experimentally analyzed the performance of newly proposed different variants. We demonstrate improved performance against the baselines and the recently proposed subspace learning methods for one-class classification. Fahad Sohrab, Alexandros Iosifidis, Moncef Gabbouj, Jenni Raitoharju |
Pattern Recognit. | 3 |
| 2023 | MAMIQA: No-Reference Image Quality Assessment Based on Multiscale Attention Mechanism With Natural Scene StatisticsabstractNo-Reference Image Quality Assessment aims to evaluate the perceptual quality of an image, according to human perception. Many recent studies use Transformers to assign different self-attention mechanisms to distinguish regions of an image, simulating the perception of the human visual system (HVS). However, the quadratic computational complexity caused by the self-attention mechanism is time-consuming and expensive. Meanwhile, the image resizing in the feature extraction stage loses the full-size image quality. To address these issues, we propose a lightweight attention mechanism using decomposed large-kernel convolutions to extract multiscale features, and a novel feature enhancement module to simulate HVS. We also propose to compensate the information loss caused by image resizing, with supplementary features from natural scene statistics. Experimental results on five standard datasets show that the proposed method surpasses the SOTA, while significantly reducing the computational costs. Li Yu 0004, Farhad Pakdaman, Miaogen Ling, Moncef Gabbouj |
IEEE Signal Process. Lett. | 5 |
| 2023 | Non-Local Color Compensation Network for Intrinsic Image DecompositionabstractSingle image-based intrinsic image decomposition attempts to separate one input image into several intrinsic components, which is inherently an under-constrained problem. Some recent works have been proposed to estimate the intrinsic components using encoder-decoder structures. However, they generally lack exploration of the different component-oriented feature constraints and feature selection processes. In this paper, a non-local color compensation network (NCCNet) is proposed. Firstly, the hue and value channels of HSV color space are used as the complementary information for RGB images for the estimation of albedo and shading, respectively. The color space representation serves as an external constraint, which does not require expensive sensors or complicated computations. Secondly, an integrated non-local attention scheme is proposed to describe the relations of non-adjacent regions with a lower computational complexity compared to traditional methods. Then the non-local and local attention are combined to describe correlations among features and used as feature selectors between the encoder and decoder. Thirdly, the mutual constraint between albedo and shading is also explored in the network to further optimize the process. In order to train the network, a unified mutual exclusion loss function is proposed. Extensive experiments are conducted on several popular datasets, and the proposed NCCNet achieves improved performance with comparable computational cost compared to competing methods. Xiaoyue Jiang, Zhaoqiang Xia, Moncef Gabbouj, Jinye Peng 0001, Xiaoyi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Generalized Tensor Summation Compressive Sensing Network (GTSNET): An Easy to Learn Compressive Sensing OperationabstractThe efforts in compressive sensing (CS) literature can be divided into two groups: finding a measurement matrix that preserves the compressed information at its maximum level, and finding a robust reconstruction algorithm. In the traditional CS setup, the measurement matrices are selected as random matrices, and optimization-based iterative solutions are used to recover the signals. Using random matrices when handling large or multi-dimensional signals is cumbersome especially when it comes to iterative optimizations. Recent deep learning-based solutions increase reconstruction accuracy while speeding up recovery, but jointly learning the whole measurement matrix remains challenging. For this reason, state-of-the-art deep learning CS solutions such as convolutional compressive sensing network (CSNET) use block-wise CS schemes to facilitate learning. In this work, we introduce a separable multi-linear learning of the CS matrix by representing the measurement signal as the summation of the arbitrary number of tensors. As compared to block-wise CS, tensorial learning eases blocking artifacts and improves performance, especially at low measurement rates (MRs), such as [Formula: see text]. The software implementation of the proposed network is publicly shared at https://github.com/mehmetyamac/GTSNET. Mehmet Yamac, Ugur Akpinar, Erdem Sahin, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Image Process. | 5 |
| 2023 | Robust Peak Detection for Holter ECGs by Self-Organized Operational Neural NetworksabstractAlthough numerous R-peak detectors have been proposed in the literature, their robustness and performance levels may significantly deteriorate in low-quality and noisy signals acquired from mobile electrocardiogram (ECG) sensors, such as Holter monitors. Recently, this issue has been addressed by deep 1-D convolutional neural networks (CNNs) that have achieved state-of-the-art performance levels in Holter monitors; however, they pose a high complexity level that requires special parallelized hardware setup for real-time processing. On the other hand, their performance deteriorates when a compact network configuration is used instead. This is an expected outcome as recent studies have demonstrated that the learning performance of CNNs is limited due to their strictly homogenous configuration with the sole linear neuron model. This has been addressed by operational neural networks (ONNs) with their heterogenous network configuration encapsulating neurons with various nonlinear operators. In this study, to further boost the peak detection performance along with an elegant computational efficiency, we propose 1-D Self-Organized ONNs (Self-ONNs) with generative neurons. The most crucial advantage of 1-D Self-ONNs over the ONNs is their self-organization capability that voids the need to search for the best operator set per neuron since each generative neuron has the ability to create the optimal operator during training. The experimental results over the China Physiological Signal Challenge-2020 (CPSC) dataset with more than one million ECG beats show that the proposed 1-D Self-ONNs can significantly surpass the state-of-the-art deep CNN with less computational complexity. Results demonstrate that the proposed solution achieves a 99.10% F1-score, 99.79% sensitivity, and 98.42% positive predictivity in the CPSC dataset, which is the best R-peak detection performance ever achieved. Moncef Gabbouj, Serkan Kiranyaz, Junaid Malik, Muhammad Uzair Zahid, Turker Ince, Muhammad E. H. Chowdhury, Amith Khandakar, Anas M. Tahir |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Convolutional Sparse Support Estimator Network (CSEN): From Energy-Efficient Support Estimation to Learning-Aided Compressive SensingabstractSupport estimation (SE) of a sparse signal refers to finding the location indices of the nonzero elements in a sparse representation. Most of the traditional approaches dealing with SE problems are iterative algorithms based on greedy methods or optimization techniques. Indeed, a vast majority of them use sparse signal recovery (SR) techniques to obtain support sets instead of directly mapping the nonzero locations from denser measurements (e.g., compressively sensed measurements). This study proposes a novel approach for learning such a mapping from a training set. To accomplish this objective, the convolutional sparse support estimator networks (CSENs), each with a compact configuration, are designed. The proposed CSEN can be a crucial tool for the following scenarios: 1) real-time and low-cost SE can be applied in any mobile and low-power edge device for anomaly localization, simultaneous face recognition, and so on and 2) CSEN's output can directly be used as "prior information," which improves the performance of sparse SR algorithms. The results over the benchmark datasets show that state-of-the-art performance levels can be achieved by the proposed approach with a significantly reduced computational complexity. Mehmet Yamac, Mete Ahishali, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | SRL-SOA: Self-Representation Learning with Sparse 1D-Operational Autoencoder for Hyperspectral Image Band SelectionabstractThe band selection in the hyperspectral image (HSI) data processing is an important task considering its effect on the computational complexity and accuracy. In this work, we propose a novel framework for the band selection problem: Self-Representation Learning (SRL) with Sparse 1D-Operational Autoencoder (SOA). The proposed SLR-SOA approach introduces a novel autoencoder model, SOA, that is designed to learn a representation domain where the data are sparsely represented. Moreover, the network composes of 1D-operational layers with the non-linear neuron model. Hence, the learning capability of neurons (filters) is greatly improved with shallow architectures. Using compact architectures is especially crucial in autoencoders as they tend to overfit easily because of their identity mapping objective. Overall, we show that the proposed SRL-SOA band selection approach outperforms the competing methods over two HSI data including Indian Pines and Salinas-A considering the achieved land cover classification accuracies. The software implementation of the SRL-SOA approach is shared publicly1. Mete Ahishali, Serkan Kiranyaz, Iftikhar Ahmad 0001, Moncef Gabbouj |
ICIP | 4 |
| 2022 | Osegnet: Operational Segmentation Network for Covid-19 Detection Using Chest X-Ray ImagesabstractCoronavirus disease 2019 (COVID-19) has been diagnosed automatically using Machine Learning algorithms over chest X-ray (CXR) images. However, most of the earlier studies used Deep Learning models over scarce datasets bearing the risk of overfitting. Additionally, previous studies have revealed the fact that deep networks are not reliable for classification since their decisions may originate from irrelevant areas on the CXRs. Therefore, in this study, we propose Operational Segmentation Network (OSegNet) that performs detection by segmenting COVID-19 pneumonia for a reliable diagnosis. To address the data scarcity encountered in training and especially in evaluation, this study extends the largest COVID-19 CXR dataset: QaTa-COV19 with 121,378 CXRs including 9258 COVID-19 samples with their corresponding ground-truth segmentation masks that are publicly shared with the research community. Consequently, OSegNet has achieved a detection performance with the highest accuracy of 99.65% among the state-of-the-art deep models with 98.09% precision. Aysen Degerli, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
ICIP | 4 |
| 2022 | Region-Of-Interest Coding Schemes For Http Adaptive Streaming With VVCabstractIndependently coding areas in a video have potential for improved coding efficiency in Region-of-Interest (RoI) streaming applications where clients can freely switch between RoI and full view streams. This paper investigates three schemes for RoI coding in streaming scenarios based on VVC to enable harnessing open GOP coding efficiency gains. Two schemes are based on Motion-Constrained-Tile-Sets while a third scheme relies on the new VVC feature referred to as independent subpictures. The reported results show substantial bitrate gains for the RoI streams compared to naïve closed-GOP coding of up to -8.68% YUV BD-rate while for the full stream minor coding efficiency losses are reported. Robert Skupin, Yago Sánchez de la Fuente, Cornelius Hellge, Thomas Schierl, Moncef Gabbouj |
ICIP | 5 |
| 2022 | Self-attention fusion for audiovisual emotion recognition with incomplete dataabstractIn this paper, we consider the problem of multi-modal data analysis with a use case of audiovisual emotion recognition. We propose an architecture capable of learning from raw data and describe three variants of it with distinct modality fusion mechanisms. While most of the previous works consider the ideal scenario of presence of both modalities at all times during inference, we evaluate the robustness of the model in the unconstrained settings where one modality is absent or noisy, and propose a method to mitigate these limitations in a form of modality dropout. Most importantly, we find that following this approach not only improves performance drastically under the absence/noisy representations of one modality, but also improves the performance in a standard ideal setting, outperforming the competing methods. Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj |
ICPR | 3 |
| 2022 | Improved Domain Adaptation Approach for Bearing Fault DiagnosisabstractApplication of domain adaptation techniques to predictive maintenance of modern electric rotating machinery (RM) has significant potential with the goal of transferring or adaptation of a fault diagnosis model developed for one machine to be generalized on new machines and/or new working conditions. The generalized nonlinear extension of conventional convolutional neural networks (CNNs), the self-organized operational neural networks (Self-ONNs) are known to enhance the learning capability of CNN by introducing non-linear neuron models and further heterogeneity in the network configuration. In this study, first the state-of-the-art 1D CNNs and Self-ONNs are tested for cross-domain performance. Then, we propose to utilize Self-ONNs as feature extractor in the well-known domain-adversarial neural networks (DANN) to enhance its domain adaptation performance. Experimental results over the benchmark Case Western Reserve University (CWRU) real vibration data set for bearing fault diagnosis across different load domains demonstrate the effectiveness and feasibility of the proposed domain adaptation approach with similar computational complexity. Turker Ince, Sertac Kilickaya, Levent Eren, Ozer Can Devecioglu, Serkan Kiranyaz, Moncef Gabbouj |
IECON | 6 |
| 2022 | OpenDR: An Open Toolkit for Enabling High Performance, Low Footprint Deep Learning for RoboticsabstractExisting Deep Learning (DL) frameworks typically do not provide ready-to-use solutions for robotics, where very specific learning, reasoning, and embodiment problems exist. Their relatively steep learning curve and the different methodologies employed by DL compared to traditional approaches, along with the high complexity of DL models, which often leads to the need of employing specialized hardware accelerators, further increase the effort and cost needed to employ DL models in robotics. Also, most of the existing DL methods follow a static inference paradigm, as inherited by the traditional computer vision pipelines, ignoring active perception, which can be employed to actively interact with the environment in order to increase perception accuracy. In this paper, we present the Open Deep Learning Toolkit for Robotics (OpenDR). OpenDR aims at developing an open, non-proprietary, efficient, and modular toolkit that can be easily used by robotics companies and research institutions to efficiently develop and deploy AI and cognition technologies to robotics applications, providing a solid step towards addressing the aforementioned challenges. We also detail the design choices, along with an abstract interface that was created to overcome these challenges. This interface can describe various robotic tasks, spanning beyond traditional DL cognition and inference, as known by existing frameworks, incorporating openness, homogeneity and robotics-oriented perception e.g., through active perception, as its core design principles. Nikolaos Passalis, S. Pedrazzi, Robert Babuska, Wolfram Burgard, D. Dias, F. Ferro, Moncef Gabbouj, Ole Green, Alexandros Iosifidis, Erdal Kayacan, Jens Kober, O. Michel, Nikos Nikolaidis 0001, Paraskevi Nousi, Roel Pieters, Maria Tzelepi, Abhinav Valada, Anastasios Tefas |
IROS | 7 |
| 2022 | High-frequency guided CNN for video compression artifacts reductionabstractIn this paper, we propose a high-frequency guided CNN for video compression artifacts reduction. In the proposed method, high frequency component in Y channel is extracted and used to guide the quality enhancement of all Y, U, V channels. As high frequency component contains the edge and contour information of the objects in the image, which is of vital importance to both subjective and objective quality. In general, the proposed method consists of two modules: the high frequency guidance module and the quality enhancement module. The high-frequency guidance module uses multiple octave convolutions to extract the high-frequency component in Y channel and then fuse it into the features of Y, U, and V channels. While in the quality enhancement module, multiple CNN residual blocks are used for the quality enhancement of Y, U, and V channels. The proposed method was integrated into both HM-16.22 and VTM-16.0. The results on the JVET test sequence under All Intra configuration shows the effectiveness of the proposed method. Compared with HEVC, the proposed method achieves the average BD-rate reductions of -12.3%, -22.7% and -23.5% for Y, U and V channels respectively. Compared with VVC, the average BD-rate reductions are -6.7%, -12.3% and -13.2% correspondingly. Li Yu 0004, Wenshuai Chang, Moncef Gabbouj |
VCIP | 4 |
| 2022 | Remote Multilinear Compressive Learning With Adaptive CompressionabstractMultilinear compressive learning (MCL) is an efficient signal acquisition and learning paradigm for multidimensional signals. The level of signal compression affects the detection or classification performance of an MCL model, with higher compression rates often associated with lower inference accuracy. However, higher compression rates are more amenable to a wider range of applications, especially those that require low operating bandwidth and minimal energy consumption such as Internet of Things (IoT) applications. Many communication protocols provide support for adaptive data transmission to maximize the throughput and minimize energy consumption. By developing compressive sensing and learning models that can operate with an adaptive compression rate, we can maximize the informational content throughput of the whole application. In this article, we propose a novel optimization scheme that enables such a feature for MCL models. Our proposal enables the practical implementation of adaptive compressive signal acquisition and inference systems. Experimental results demonstrated that the proposed approach can significantly reduce the amount of computations required during the training phase of remote learning systems but also improve the informational content throughput via adaptive-rate sensing. Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Internet Things J. | 2 |
| 2022 | Feedforward neural networks initialization based on discriminant learningabstractIn this paper, a novel data-driven method for weight initialization of Multilayer Perceptrons and Convolutional Neural Networks based on discriminant learning is proposed. The approach relaxes some of the limitations of competing data-driven methods, including unimodality assumptions, limitations on the architectures related to limited maximal dimensionalities of the corresponding projection spaces, as well as limitations related to high computational requirements due to the need of eigendecomposition on high-dimensional data. We also consider assumptions of the method on the data and propose a way to account for them in a form of a new normalization layer. The experiments on three large-scale image datasets show improved accuracy of the trained models compared to competing random-based and data-driven weight initialization methods, as well as better convergence properties in certain cases. Kateryna Chumachenko, Alexandros Iosifidis, Moncef Gabbouj |
Neural Networks | 3 |
| 2022 | Neural texture transfer assisted video coding with adaptive up-sampling
Li Yu 0004, Wenshuai Chang, Weize Quan, Jimin Xiao, Dong-Ming Yan 0001, Moncef Gabbouj |
Signal Process. Image Commun. | 6 |
| 2022 | Saliency-Based Multilabel Linear Discriminant AnalysisabstractLinear discriminant analysis (LDA) is a classical statistical machine-learning method, which aims to find a linear data transformation increasing class discrimination in an optimal discriminant subspace. Traditional LDA sets assumptions related to the Gaussian class distributions and single-label data annotations. In this article, we propose a new variant of LDA to be used in multilabel classification tasks for dimensionality reduction on original data to enhance the subsequent performance of any multilabel classifier. A probabilistic class saliency estimation approach is introduced for computing saliency-based weights for all instances. We use the weights to redefine the between-class and within-class scatter matrices needed for calculating the projection matrix. We formulate six different variants of the proposed saliency-based multilabel LDA (SMLDA) based on different prior information on the importance of each instance for their class(es) extracted from labels and features. Our experiments show that the proposed SMLDA leads to performance improvements in various multilabel classification problems compared to several competing dimensionality reduction methods. Lei Xu 0036, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
IEEE Trans. Cybern. | 4 |
| 2022 | Hierarchical Symbol Transition Entropy: A Novel Feature Extractor for Machinery Health MonitoringabstractThis article develops a novel collaborative health monitoring framework based on hierarchical symbol transition entropy (HSTE) and 2-D-extreme learning machine (2-D-ELM) without fusion. In the proposed framework, a novel metric called symbol transition entropy (STE) is first presented to evaluate the dynamical complexity of time series through multistep transition and the joint probability distribution of the symbol state and its transition state. Compared with the existing entropy algorithms, STE has better robustness and captures more detailed dynamical changes. Subsequently, a new feature representation method called HSTE is proposed by combining STE with the hierarchical analysis. The two-order tensor features can be constructed for multichannel data by stacking HSTE values extracted from each single-channel data. Finally, 2-D-ELM is incorporated to identify the extracted two-order tensor features without vectorization. The feasibility of the proposed schemes is verified through simulation and experimental studies, and the final results confirm that the developed schemes have better performance than the existing entropy-based collaborative fault diagnosis methods. Moncef Gabbouj, Minping Jia, Zhinong Li |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Learning to ignore: rethinking attention in CNNs
Firas Laakom, Kateryna Chumachenko, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
BMVC | 5 |
| 2021 | Robust channel-wise illumination estimation
Firas Laakom, Jenni Raitoharju, Jarno Nikkanen, Alexandros Iosifidis, Moncef Gabbouj |
BMVC | 5 |
| 2021 | Ensembling Object Detectors for Image and Video Data AnalysisabstractIn this paper, we propose a method for ensembling the outputs of multiple object detectors for improving detection performance and precision of bounding boxes on image data. We further extend it to video data by proposing a two-stage tracking-based scheme for detection refinement. The proposed method can be used as a standalone approach for improving object detection performance, or as a part of a framework for faster bounding box annotation in unseen datasets, assuming that the objects of interest are those present in some common public datasets. Kateryna Chumachenko, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
ICASSP | 4 |
| 2021 | Mind The Structure: Adopting Structural Information For Deep Neural Network CompressionabstractDeep neural networks have huge number of parameters and require large number of bits for representation. This hinders their adoption in decentralized environments where model transfer among different parties is a characteristic of the environment while the communication bandwidth is limited. Parameter quantization is a compression approach to address this challenge by reducing the number of bits required to represent a model, e.g. a neural network. However, majority of existing neural network quantization methods do not exploit structural information of layers and parameters during quantization. In this paper, focusing on Convolutional Neural Networks (CNNs), we present a novel quantization approach by employing the structural information of neural network layers and their corresponding parameters. Starting from a pre-trained CNN, we categorize network parameters into different groups based on the similarity of their layers and their spatial structure. Parameters of each group are independently clustered and the centroid of each cluster is used as representative for all parameters in the cluster. Finally, the centroids and the cluster indexes of the parameters are used as a compact representation of the parameters. Experiments with two different tasks, i.e., acoustic scene classification and image compression, demonstrate the effectiveness of the proposed approach. Homayun Afrabandpey, Anton Muravev, Hamed Rezazadegan Tavakoli, Honglei Zhang 0001, Francesco Cricri, Moncef Gabbouj, Emre Aksu |
ICIP | 6 |
| 2021 | Reliable Covid-19 Detection using Chest X-Ray ImagesabstractCoronavirus disease 2019 (COVID-19) has emerged the need for computer-aided diagnosis with automatic, accurate, and fast algorithms. Recent studies have applied Machine Learning algorithms for COVID-19 diagnosis over chest X-ray (CXR) images. However, the data scarcity in these studies prevents a reliable evaluation with the potential of overfitting and limits the performance of deep networks. Moreover, these networks can discriminate COVID-19 pneumonia usually from healthy subjects only or occasionally, from limited pneumonia types. Thus, there is a need for a robust and accurate COVID-19 detector evaluated over a large CXR dataset. To address this need, in this study, we propose a reliable COVID-19 detection network: ReCovNet, which can discriminate COVID-19 pneumonia from 14 different thoracic diseases and healthy subjects. To accomplish this, we have compiled the largest COVID-19 CXR dataset: QaTa-COV19 with 124,616 images including 4603 COVID-19 samples. The proposed ReCovNet achieved a detection performance with 98.57% sensitivity and 99.77% specificity. Aysen Degerli, Mete Ahishali, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
ICIP | 5 |
| 2021 | Bm3d Vs 2-Layer OnnabstractDespite their recent success on image denoising, the need for deep and complex architectures still hinders the practical usage of CNNs. Older but computationally more efficient methods such as BM3D remain a popular choice, especially in resource-constrained scenarios. In this study, we aim to find out whether compact neural networks can learn to produce competitive results as compared to BM3D for AWGN image denoising. To this end, we conFigure networks with only two hidden layers and employ different neuron models and layer widths for comparing the performance with BM3D across different AWGN noise levels. Our results conclusively show that the recently proposed self-organized variant of operational neural networks based on a generative neuron model (Self-ONNs) is not only a better choice as compared to CNNs, but also provide competitive results as compared to BM3D and even significantly surpass it for high noise levels. Junaid Malik, Serkan Kiranyaz, Mehmet Yamac, Moncef Gabbouj |
ICIP | 4 |
| 2021 | Monte Carlo Dropout Ensembles for Robust Illumination EstimationabstractComputational color constancy is a preprocessing step used in many camera systems. The main aim is to discount the effect of the illumination on the colors in the scene and restore the original colors of the objects. Recently, several deep learning-based approaches have been proposed to solve this problem and they often led to state-of-the-art performance in terms of average errors. However, for extreme samples, these methods fail and lead to high errors. In this paper, we address this limitation by proposing to aggregate different deep learning methods according to their output uncertainty. We estimate the relative uncertainty of each approach using Monte Carlo dropout and the final illumination estimate is obtained as the sum of the different model estimates weighted by the log-inverse of their corresponding uncertainties. The proposed framework leads to state-of-the-art performance on INTEL-TAU dataset. Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Jarno Nikkanen, Moncef Gabbouj |
IJCNN | 5 |
| 2021 | Speech Command Recognition in Computationally Constrained Environments with a Quadratic Self-Organized Operational LayerabstractAutomatic classification of speech commands has revolutionized human computer interactions in robotic applications. However, employed recognition models usually follow the methodology of deep learning with complicated networks which are memory and energy hungry. So, there is a need to either squeeze these complicated models or use more efficient lightweight models in order to be able to implement the resulting classifiers on embedded devices. In this paper, we pick the second approach and propose a network layer to enhance the speech command recognition capability of a lightweight network and demonstrate the result via experiments. The employed method borrows the ideas of Taylor expansion and quadratic forms to construct a better representation of features in both input and hidden layers. This richer representation results in recognition accuracy improvement as shown by extensive experiments on Google speech commands (GSC) and synthetic speech commands (SSC) datasets. Mohammad Soltanian, Junaid Malik, Jenni Raitoharju, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
IJCNN | 6 |
| 2021 | Content-adaptive convolutional neural network post-processing filterabstractNeural Network (NN)-based coding techniques are being developed for hybrid video coding schemes, such as the Versatile Video Coding (VVC) standard. In-loop filters and postprocessing filters are two types of coding tools that aim to improve the visual quality of the reconstructed content. These tools are usually trained on large video or image datasets with varying content, but they are rarely adaptive to different content types. This problem is addressed with the proposed content-adaptive Convolutional Neural Network (CNN) post-processing filter. The proposed approach is content-adaptive in two ways. Firstly, a relatively simple CNN is pre-trained on a general video dataset and then fine-tuned on the video to be coded. Since only the bias terms of the CNN are fine-tuned, the signalling overhead is reduced. Secondly, a scaling factor indicates the influence of the CNN post-processing filter on the final reconstruction. The CNN post-processing filter is evaluated on top of VVC Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0 and, overall, it can save 2.37% (Y), 3.63% (U), 2.24% (V) Bjøntegaard Delta rate (BD-rate) in the Random Access (RA) configuration. María Santamaría 0001, Yat-Hong Lam, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela, Esa Rahtu, Moncef Gabbouj |
ISM | 9 |
| 2021 | AnomalyHop: An SSL-based Image Anomaly Localization MethodabstractAn image anomaly localization method based on the successive subspace learning (SSL) framework, called Anomaly-Hop, is proposed in this work. AnomalyHop consists of three modules: 1) feature extraction via successive subspace learning (SSL), 2) normality feature distributions modeling via Gaussian models, and 3) anomaly map generation and fusion. Comparing with state-of-the-art image anomaly localization methods based on deep neural networks (DNNs), AnomalyHop is mathematically transparent, easy to train, and fast in its inference speed. Besides, its area under the ROC curve (ROC-AUC) performance on the MVTec AD dataset is 95.9%, which is among the best of several benchmarking methods. Kaitai Zhang, Bin Wang 0040, Wei Wang 0352, Fahad Sohrab, Moncef Gabbouj, C.-C. Jay Kuo |
VCIP | 5 |
| 2021 | Exploiting heterogeneity in operational neural networks by synaptic plasticityabstractAbstract The recently proposed network model, Operational Neural Networks (ONNs), can generalize the conventional Convolutional Neural Networks (CNNs) that are homogenous only with a linear neuron model. As a heterogenous network model, ONNs are based on a generalized neuron model that can encapsulateanyset of non-linear operators to boost diversity and to learn highly complex and multi-modal functions or spaces with minimal network complexity and training data. However, the default search method to find optimal operators in ONNs, the so-called Greedy Iterative Search (GIS) method, usually takes several training sessions to find a single operator set per layer. This is not only computationally demanding, also the network heterogeneity is limited since the same set of operators will then be used for all neurons in each layer. To address this deficiency and exploit a superior level of heterogeneity, in this study the focus is drawn on searching the best-possible operator set(s) for the hidden neurons of the network based on the “Synaptic Plasticity” paradigm that poses the essential learning theory in biological neurons. During training, each operator set in the library can be evaluated by their synaptic plasticity level, ranked from the worst to the best, and an “elite” ONN can then be configured using the top-ranked operator sets found at each hidden layer. Experimental results over highly challenging problems demonstrate that the elite ONNs even with few neurons and layers can achieve a superior learning performance than GIS-based ONNs and as a result, the performance gap over the CNNs further widens. Serkan Kiranyaz, Junaid Malik, Habib Ben Abdallah, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Comput. Appl. | 6 |
| 2021 | Self-organized Operational Neural Networks with Generative NeuronsabstractOperational Neural Networks (ONNs) have recently been proposed to address the well-known limitations and drawbacks of conventional Convolutional Neural Networks (CNNs) such as network homogeneity with the sole linear neuron model. ONNs are heterogeneous networks with a generalized neuron model. However the operator search method in ONNs is not only computationally demanding, but the network heterogeneity is also limited since the same set of operators will then be used for all neurons in each layer. Moreover, the performance of ONNs directly depends on the operator set library used, which introduces a certain risk of performance degradation especially when the optimal operator set required for a particular task is missing from the library. In order to address these issues and achieve an ultimate heterogeneity level to boost the network diversity along with computational efficiency, in this study we propose Self-organized ONNs (Self-ONNs) with generative neurons that can adapt (optimize) the nodal operator of each connection during the training process. Moreover, this ability voids the need of having a fixed operator set library and the prior operator search within the library in order to find the best possible set of operators. We further formulate the training method to back-propagate the error through the operational layers of Self-ONNs. Experimental results over four challenging problems demonstrate the superior learning capability and computational efficiency of Self-ONNs over conventional ONNs and CNNs. Serkan Kiranyaz, Junaid Malik, Habib Ben Abdallah, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Networks | 6 |
| 2021 | Self-organized operational neural networks for severe image restoration problemsabstractDiscriminative learning based on convolutional neural networks (CNNs) aims to perform image restoration by learning from training examples of noisy-clean image pairs. It has become the go-to methodology for tackling image restoration and has outperformed the traditional non-local class of methods. However, the top-performing networks are generally composed of many convolutional layers and hundreds of neurons, with trainable parameters in excess of several million. We claim that this is due to the inherently linear nature of convolution-based transformation, which is inadequate for handling severe restoration problems. Recently, a non-linear generalization of CNNs, called the operational neural networks (ONN), has been shown to outperform CNN on AWGN denoising. However, its formulation is burdened by a fixed collection of well-known non-linear operators and an exhaustive search to find the best possible configuration for a given architecture, whose efficacy is further limited by a fixed output layer operator assignment. In this study, we leverage the Taylor series-based function approximation to propose a self-organizing variant of ONNs, Self-ONNs, for image restoration, which synthesizes novel nodal transformations on-the-fly as part of the learning process, thus eliminating the need for redundant training runs for operator search. In addition, it enables a finer level of operator heterogeneity by diversifying individual connections of the receptive fields and weights. We perform a series of extensive ablation experiments across three severe image restoration tasks. Even when a strict equivalence of learnable parameters is imposed, Self-ONNs surpass CNNs by a considerable margin across all problems, improving the generalization performance by up to 3 dB in terms of PSNR. Junaid Malik, Serkan Kiranyaz, Moncef Gabbouj |
Neural Networks | 3 |
| 2021 | Speed-up and multi-view extensions to subclass discriminant analysisabstractIn this paper, we propose a speed-up approach for subclass discriminant analysis and formulate a novel efficient multi-view solution to it. The speed-up approach is developed based on graph embedding and spectral regression approaches that involve eigendecomposition of the corresponding Laplacian matrix and regression to its eigenvectors. We show that by exploiting the structure of the between-class Laplacian matrix, the eigendecomposition step can be substituted with a much faster process. Furthermore, we formulate a novel criterion for multi-view subclass discriminant analysis and show that an efficient solution to it can be obtained in a similar manner to the single-view case. We evaluate the proposed methods on nine single-view and nine multi-view datasets and compare them with related existing approaches. Experimental results show that the proposed solutions achieve competitive performance, often outperforming the existing methods. At the same time, they significantly decrease the training time. Kateryna Chumachenko, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 4 |
| 2021 | Multimodal subspace support vector data descriptionabstractIn this paper, we propose a novel method for projecting data from multiple modalities to a new subspace optimized for one-class classification. The proposed method iteratively transforms the data from the original feature space of each modality to a new common feature space along with finding a joint compact description of data coming from all the modalities. For data in each modality, we define a separate transformation to map the data from the corresponding feature space to the new optimized subspace by exploiting the available information from the class of interest only. We also propose different regularization strategies for the proposed method and provide both linear and non-linear formulations. The proposed Multimodal Subspace Support Vector Data Description outperforms all the competing methods using data from a single modality or fusing data from all modalities in four out of five datasets. Fahad Sohrab, Jenni Raitoharju, Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 4 |
| 2021 | SVM based approach for complexity control of HEVC intra coding
Farhad Pakdaman, Li Yu 0004, Mahmoud Reza Hashemi, Mohammed Ghanbari 0001, Moncef Gabbouj |
Signal Process. Image Commun. | 5 |
| 2021 | CasQNet: Intrinsic Image Decomposition Based on Cascaded Quotient NetworkabstractIntrinsic image analysis plays an important role for image understanding, since it can provide accurate reflectance, shape and illumination information of the scene. However, intrinsic image analysis is an ill-posed problem which need to apply extra constrains for the decomposition of reflectance image and shading image from a single image. Recently deep neural networks are introduced for intrinsic image analysis, which can produce two intrinsic components simultaneously. In fact, the mutually exclusive relationship between reflectance image and shading image is not only a constraint for decomposition but also can improve the decomposition results. However, this relationship is always omitted in the current networks. In order to address this problem, we propose a novel deep network called as Cascaded Quotient Network (CasQNet) for intrinsic image decomposition. The CasQNet consists of two sub-networks: a Pyramid Mini-U-Net (PyNet) that specifically extracts the reflectance image in multi-scale and a Shading Optimization Network (SoNet) that optimizes the resulting shading. These two sub-networks are cascaded by a quotient operation, which directly enforces the mutually exclusive relationship between reflectance image and shading image in the network architecture. In PyNet, the task of reconstructing reflectance image is achieved by a series of nested multi-scale U-Nets, which simplified the learning task for each U-Net. SoNet is designed to address the unsmooth and blur problems of extreme points caused by the quotient operation. PyNet and SoNet are trained alternately and finally jointed in cascaded structure. Furthermore, we combine multiple loss functions, which consist of data loss, correlation loss and reconstruction loss, for improving the learning effectiveness. To evaluate our proposed algorithm, extensive experiments are performed on three datasets, i.e., ShapeNet, BOLD Surface and MIT Intrinsic Image datasets. Qualitative and quantitative results show that our model achieves the best performance compared to the state-of-the-art methods. Yupeng Ma, Xiaoyue Jiang, Zhaoqiang Xia, Moncef Gabbouj, Xiaoyi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Multi-Level Reversible Data Anonymization via Compressive Sensing and Data HidingabstractRecent advances in intelligent surveillance systems have enabled a new era of smart monitoring in a wide range of applications from health monitoring to homeland security. However, this boom in data gathering, analyzing and sharing brings in also significant privacy concerns. We propose a Compressive Sensing (CS) based data encryption that is capable of both obfuscating selected sensitive parts of documents and compressively sampling, hence encrypting both sensitive and non-sensitive parts of the document. The scheme uses a data hiding technique on CS-encrypted signal to preserve the one-time use obfuscation matrix. The proposed privacy-preserving approach offers a low-cost multi-tier encryption system that provides different levels of reconstruction quality for different classes of users, e.g., semi-authorized, full-authorized. As a case study, we develop a secure video surveillance system and analyze its performance. Mehmet Yamac, Mete Ahishali, Nikolaos Passalis, Jenni Raitoharju, Bülent Sankur, Moncef Gabbouj |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2021 | Deep Multi-View Learning to RankabstractWe study the problem of learning to rank from multiple information sources. Though multi-view learning and learning to rank have been studied extensively leading to a wide range of applications, multi-view learning to rank as a synergy of both topics has received little attention. The aim of the paper is to propose a composite ranking method while keeping a close correlation with the individual rankings simultaneously. We present a generic framework for multi-view subspace learning to rank (MvSL2R), and two novel solutions are introduced under the framework. The first solution captures information of feature mappings from within each view as well as across views using autoencoder-like networks. Novel feature embedding methods are formulated in the optimization of multi-view unsupervised and discriminant autoencoders. Moreover, we introduce an end-to-end solution to learning towards both the joint ranking objective and the individual rankings. The proposed solution enhances the joint ranking with minimum view-specific ranking loss, so that it can achieve the maximum global view agreements in a single optimization process. The proposed method is evaluated on three different ranking problems, i.e., university ranking, multi-view lingual text ranking, and image data ranking, providing superior results compared to related methods. Guanqun Cao, Alexandros Iosifidis, Moncef Gabbouj, Vijay Raghavan 0001, Raju N. Gottumukkala |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Hypersphere-Based Weight Imprinting for Few-Shot Learning on Embedded DevicesabstractWeight imprinting (WI) was recently introduced as a way to perform gradient descent-free few-shot learning. Due to this, WI was almost immediately adapted for performing few-shot learning on embedded neural network accelerators that do not support back-propagation, e.g., edge tensor processing units. However, WI suffers from many limitations, e.g., it cannot handle novel categories with multimodal distributions and special care should be given to avoid overfitting the learned embeddings on the training classes since this can have a devastating effect on classification accuracy (for the novel categories). In this article, we propose a novel hypersphere-based WI approach that is capable of training neural networks in a regularized, imprinting-aware way effectively overcoming the aforementioned limitations. The effectiveness of the proposed method is demonstrated using extensive experiments on three image data sets. Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Multilinear Compressive LearningabstractCompressive learning (CL) is an emerging topic that combines signal acquisition via compressive sensing (CS) and machine learning to perform inference tasks directly on a small number of measurements. Many data modalities naturally have a multidimensional or tensorial format, with each dimension or tensor mode representing different features such as the spatial and temporal information in video sequences or the spatial and spectral information in hyperspectral images. However, in existing CL frameworks, the CS component utilizes either random or learned linear projection on the vectorized signal to perform signal acquisition, thus discarding the multidimensional structure of the signals. In this article, we propose multilinear CL (MCL), a framework that takes into account the tensorial nature of multidimensional signals in the acquisition step and builds the subsequent inference model on the structurally sensed measurements. Our theoretical complexity analysis shows that the proposed framework is more efficient compared to its vector-based counterpart in both memory and computation requirement. With extensive experiments, we also empirically show that our MCL framework outperforms the vector-based framework in object classification and face recognition tasks, and scales favorably when the dimensionalities of the original signals increase, making it highly efficient for high-dimensional multidimensional signals. Dat Thanh Tran, Mehmet Yamac, Aysen Degerli, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Convolutional Sparse Support Estimator-Based COVID-19 Recognition From X-Ray ImagesabstractCoronavirus disease (COVID-19) has been the main agenda of the whole world ever since it came into sight. X-ray imaging is a common and easily accessible tool that has great potential for COVID-19 diagnosis and prognosis. Deep learning techniques can generally provide state-of-the-art performance in many classification tasks when trained properly over large data sets. However, data scarcity can be a crucial obstacle when using them for COVID-19 detection. Alternative approaches such as representation-based classification [collaborative or sparse representation (SR)] might provide satisfactory performance with limited size data sets, but they generally fall short in performance or speed compared to the neural network (NN)-based methods. To address this deficiency, convolution support estimation network (CSEN) has recently been proposed as a bridge between representation-based and NN approaches by providing a noniterative real-time mapping from query sample to ideally SR coefficient support, which is critical information for class decision in representation-based techniques. The main premises of this study can be summarized as follows: 1) A benchmark X-ray data set, namely QaTa-Cov19, containing over 6200 X-ray images is created. The data set covering 462 X-ray images from COVID-19 patients along with three other classes; bacterial pneumonia, viral pneumonia, and normal. 2) The proposed CSEN-based classification scheme equipped with feature extraction from state-of-the-art deep NN solution for X-ray images, CheXNet, achieves over 98% sensitivity and over 95% specificity for COVID-19 recognition directly from raw X-ray images when the average performance of 5-fold cross validation over QaTa-Cov19 data set is calculated. 3) Having such an elegant COVID-19 assistive diagnosis performance, this study further provides evidence that COVID-19 induces a unique pattern in X-rays that can be discriminated with high accuracy. Mehmet Yamac, Mete Ahishali, Aysen Degerli, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Adaptive Normalization for Forecasting Limit Order Book Data Using Convolutional Neural NetworksabstractDeep learning models are capable of achieving state-of-the-art performance on a wide range of time series analysis tasks. However, their performance crucially depends on the employed normalization scheme, while they are usually unable to efficiently handle non-stationary features without first appropriately pre-processing them. These limitations impact the performance of deep learning models, especially when used for forecasting financial time series, due to their non-stationary and multimodal nature. In this paper we propose a data-driven adaptive normalization layer which is capable of learning the most appropriate normalization scheme that should be applied on the data. To this end, the proposed method first identifies the distribution from which the data were generated and then it dynamically shifts and scales them in order to facilitate the task at hand. The proposed nor-malization scheme is fully differentiable and it is trained in an end-to-end fashion along with the rest of the parameters of the model. The proposed method leads to significant performance improvements over several competitive normalization approaches, as demonstrated using a large-scale limit order book dataset. Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
ICASSP | 4 |
| 2020 | Incremental Fast Subclass Discriminant AnalysisabstractThis paper proposes an incremental solution to Fast Subclass Discriminant Analysis (fastSDA). We present an exact and an approximate linear solution, along with an approximate kernelized variant. Extensive experiments on eight image datasets with different incremental batch sizes show the superiority of the proposed approach in terms of training time and accuracy being equal or close to fastSDA solution and outperforming other methods. Kateryna Chumachenko, Jenni Raitoharju, Moncef Gabbouj, Alexandros Iosifidis |
ICIP | 3 |
| 2020 | Probabilistic Color ConstancyabstractIn this paper, we propose a novel unsupervised color constancy method, called Probabilistic Color Constancy (PCC). We define a framework for estimating the illumination of a scene by weighting the contribution of different image regions using a graph-based representation of the image. To estimate the weight of each (super-)pixel, we rely on two assumptions: (Super-)pixels with similar colors contribute similarly and darker (super-)pixels contribute less. The resulting system has one global optimum solution. The proposed method achieves competitive performance, compared to the state-of-the-art, on INTEL-TAU dataset. Firas Laakom, Jenni Raitoharju, Alexandros Iosifidis, Uygar Tuna, Jarno Nikkanen, Moncef Gabbouj |
ICIP | 6 |
| 2020 | Complexity Analysis Of Next-Generation VVC Encoding And DecodingabstractWhile the next generation video compression standard, Versatile Video Coding (VVC), provides a superior compression efficiency, its computational complexity dramatically increases. This paper thoroughly analyzes this complexity for both encoder and decoder of VVC Test Model 6, by quantifying the complexity break-down for each coding tool and measuring the complexity and memory requirements for VVC encoding/decoding. These extensive analyses are performed for six video sequences of 720p, 1080p, and 2160p, under Low-Delay (LD), Random-Access (RA), and All-Intra (AI) conditions (a total of 320 encoding/decoding). Results indicate that the VVC encoder and decoder are 5× and 1.5× more complex compared to HEVC in LD, and 31× and 1.8× in AI, respectively. Detailed analysis of coding tools reveals that in LD on average, motion estimation tools with 53%, transformation and quantization with 22%, and entropy coding with 7% dominate the encoding complexity. In decoding, loop filters with 30%, motion compensation with 20%, and entropy decoding with 16%, are the most complex modules. Moreover, the required memory bandwidth for VVC encoding/decoding are measured through memory profiling, which are 30× and 3× of HEVC. The reported results and insights are a guide for future research and implementations of energy-efficient VVC encoder/decoder. Farhad Pakdaman, Mohammad Ali Adelimanesh, Moncef Gabbouj, Mahmoud Reza Hashemi |
ICIP | 3 |
| 2020 | Subset Sampling for Progressive Neural Network LearningabstractProgressive Neural Network Learning is a class of algorithms that incrementally construct the network's topology and optimize its parameters based on the training data. While this approach exempts the users from the manual task of designing and validating multiple network topologies, it often requires an enormous number of computations. In this paper, we propose to speed up this process by exploiting subsets of training data at each incremental training step. Three different sampling strategies for selecting the training samples according to different criteria are proposed and evaluated. We also propose to perform online hyperparameter selection during the network progression, which further reduces the overall training time. Experimental results in object, scene and face recognition problems demonstrate that the proposed approach speeds up the optimization procedure considerably while operating on par with the baseline approach exploiting the entire training set throughout the training process. Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis |
ICIP | 2 |
| 2020 | Not all domains are equally complex: Adaptive Multi-Domain LearningabstractDeep learning approaches are highly specialized and require training separate models for different tasks. Multidomain learning looks at ways to learn a multitude of different tasks, each coming from a different domain, at once. The most common approach in multi-domain learning is to form a domain agnostic model, the parameters of which are shared among all domains, and learn a small number of extra domain-specific parameters for each individual new domain. However, different domains come with different levels of difficulty; parameterizing the models of all domains using an augmented version of the domain agnostic model leads to unnecessarily inefficient solutions, especially for easy to solve tasks. We propose an adaptive parameterization approach to deep neural networks for multidomain learning. The proposed approach performs on par with the original approach while reducing by far the number of parameters, leading to efficient multi-domain learning solutions. Ali Senhaji, Jenni Raitoharju, Moncef Gabbouj, Alexandros Iosifidis |
ICPR | 3 |
| 2020 | Data Normalization for Bilinear Structures in High-Frequency Financial Time-seriesabstractFinancial time-series analysis and forecasting have been extensively studied over the past decades, yet still remain as a very challenging research topic. Since the financial market is inherently noisy and stochastic, a majority of financial time-series of interests are non-stationary, and often obtained from different modalities. This property presents great challenges and can significantly affect the performance of the subsequent analysis/forecasting steps. Recently, the Temporal Attention augmented Bilinear Layer (TABL) has shown great performances in tackling financial forecasting problems. In this paper, by taking into account the nature of bilinear projections in TABL networks, we propose Bilinear Normalization (BiN), a simple, yet efficient normalization layer to be incorporated into TABL networks to tackle potential problems posed by non-stationarity and multimodalities in the input series. Our experiments using a large scale Limit Order Book (LOB) consisting of more than 4 million order events show that BiN-TABL outperforms TABL networks using other state-of-the-arts normalization schemes by a large margin. Dat Thanh Tran, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
ICPR | 3 |
| 2020 | Generalized Operational Classifiers for Material IdentificationabstractMaterial is one of the intrinsic features of objects, and consequently material recognition plays an important role in image understanding. The same material may have various shapes and appearance, while keeping the same physical characteristic. This brings great challenges for material recognition. Besides suitable features, a powerful classifier also can improve the overall recognition performance. Due to the limitations of classical linear neurons, used in all shallow and deep neural networks, such as CNN, we propose to apply the generalized operational neurons to construct a classifier adaptively. These generalized operational perceptrons (GOP) contain a set of linear and nonlinear neurons, and possess a structure that can be built progressively. This makes GOP classifier more compact and can easily discriminate complex classes. The experiments demonstrate that GOP networks trained on a small portion of the data (4%) can achieve comparable performances to state-of-the-arts models trained on much larger portions of the dataset. Xiaoyue Jiang, Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Xiaoyi Feng |
MMSP | 5 |
| 2020 | Efficient Adaptive Inference Leveraging Bag-of-Features-based Early ExitsabstractEarly exits provide an effective way of implementing adaptive computational graphs over deep learning models. In this way it is possible to adapt them on-the-fly to the available computational resources or even to the difficulty of each input sample, reducing the energy and computational power requirements in many embedded and mobile applications. However, performing this kind of adaptive inference also comes with several challenges, since the difficulty of each sample must be estimated and the most appropriate early exit must be selected. It is worth noting that existing approaches often lead to highly unbalanced distributions over the selected early exits, reducing the efficiency of the adaptive inference process. At the same time, only a few resources can be devoted to the aforementioned process, in order to ensure that an adequate speedup will be obtained. The main contribution of this work is to provide an easy to use and tune adaptive inference approach for early exits that can overcome some of these limitations. In this way, the proposed method allows for a) obtaining a more balanced inference distribution among the early exits, b) relying on a single and interpretable hyperparameter for tuning its behavior (ranging from faster inference to higher accuracy), and c) improving the performance of the networks (increasing the accuracy and reducing the time needed for inference). Indeed, the effectiveness of the proposed method over existing approaches is demonstrated using four different image datasets. Nikolaos Passalis, Jenni Raitoharju, Moncef Gabbouj, Anastasios Tefas |
MMSP | 3 |
| 2020 | Real-time phonocardiogram anomaly detection by adaptive 1D Convolutional Neural NetworksabstractThe heart sound signals (Phonocardiogram – PCG) enable the earliest monitoring to detect a potential cardiovascular pathology and have recently become a crucial tool as a diagnostic test in outpatient monitoring to assess heart hemodynamic status. The need for an automated and accurate anomaly detection method for PCG has thus become imminent. To determine the state-of-the-art PCG classification algorithm, 48 international teams competed in the PhysioNet (CinC) Challenge in 2016 over the largest benchmark dataset with 3126 records with the classification outputs, normal (N), abnormal (A) and unsure – too noisy (U). In this study, our aim is to push this frontier further; however, we focus deliberately on the anomaly detection problem while assuming a reasonably high Signal-to-Noise Ratio (SNR) on the records. By using 1D Convolutional Neural Networks trained with a novel data purification approach, we aim to achieve the highest detection performance and real-time processing ability with significantly lower delay and computational complexity. The experimental results over the high-quality subset of the same benchmark dataset show that the proposed approach achieves both objectives. Furthermore, our findings reveal the fact that further improvements indeed require a personalized (patient-specific) approach to avoid major drawbacks of a global PCG classification approach. Serkan Kiranyaz, Morteza Zabihi, Ali Bahrami Rad, Turker Ince, Ridha Hamila, Moncef Gabbouj |
Neurocomputing | 6 |
| 2020 | Progressive Operational Perceptrons with MemoryabstractGeneralized Operational Perceptron (GOP) was proposed to generalize the linear neuron model used in the traditional Multilayer Perceptron (MLP) by mimicking the synaptic connections of biological neurons showing nonlinear neurochemical behaviours. Previously, Progressive Operational Perceptron (POP) was proposed to train a multilayer network of GOPs which is formed layer-wise in a progressive manner. While achieving superior learning performance over other types of networks, POP has a high computational complexity. In this work, we propose POPfast, an improved variant of POP that signicantly reduces the computational complexity of POP, thus accelerating the training time of GOP networks. In addition, we also propose major architectural modications of POPfast that can augment the progressive learning process of POP by incorporating an information preserving, linear projection path from the input to the output layer at each progressive step. The proposed extensions can be interpreted as a mechanism that provides direct information extracted from the previously learned layers to the network, hence the term “memory”. This allows the network to learn deeper architectures and better data representations. An extensive set of experiments in human action, object, facial identity and scene recognition problems demonstrates that the proposed algorithms can train GOP networks much faster than POPs while achieving better performance compared to original POPs and other related algorithms. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
Neurocomputing | 3 |
| 2020 | Real-time throughput prediction for cognitive Wi-Fi networks
Muhammad Asif Khan 0001, Ridha Hamila, Nasser Al-Emadi, Serkan Kiranyaz, Moncef Gabbouj |
J. Netw. Comput. Appl. | 5 |
| 2020 | Operational neural networksabstractAbstract Feed-forward, fully connected artificial neural networks or the so-called multi-layer perceptrons are well-known universal approximators. However, their learning performance varies significantly depending on the function or the solution space that they attempt to approximate. This is mainly because of their homogenous configuration based solely on the linear neuron model. Therefore, while they learn very well those problems with a monotonous, relatively simple and linearly separable solution space, they may entirely fail to do so when the solution space is highly nonlinear and complex. Sharing the same linear neuron model with two additional constraints (local connections and weight sharing), this is also true for the conventional convolutional neural networks (CNNs) and it is, therefore, not surprising that in many challenging problems only the deep CNNs with a massive complexity and depth can achieve the required diversity and the learning performance. In order to address this drawback and also to accomplish a more generalized model over the convolutional neurons, this study proposes a novel network model, called operational neural networks (ONNs), which can be heterogeneous and encapsulate neurons with any set of operators to boost diversity and to learn highly complex and multi-modal functions or spaces with minimal network complexity and training data. Finally, the training method to back-propagate the error through the operational layers of ONNs is formulated. Experimental results over highly challenging problems demonstrate the superior learning capabilities of ONNs even with few neurons and hidden layers. Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Comput. Appl. | 4 |
| 2020 | Recurrent bag-of-features for visual information analysisabstractDeep Learning (DL) has provided powerful tools for visual information analysis. For example, Convolutional Neural Networks (CNNs) are excelling in complex and challenging image analysis tasks by extracting meaningful feature vectors with high discriminative power . However, these powerful feature vectors are crushed through the pooling layers of the network, that usually implement the pooling operation in a less sophisticated manner. This can lead to significant information loss, especially in cases where the informative content of the data is sequentially distributed over the spatial or temporal dimension, e.g., videos, which often require extracting fine-grained temporal information. A novel stateful recurrent pooling approach, that can overcome the aforementioned limitations, is proposed in this paper. The proposed method is inspired by the well-known Bag-of-Features (BoF) model, but employs a stateful trainable recurrent quantizer, instead of plain static quantization, allowing for efficiently processing sequential data and encoding both their temporal, as well as their spatial aspects. The effectiveness of the proposed Recurrent BoF model to enclose spatio-temporal information compared to other competitive methods is demonstrated using six different datasets and two different tasks. Marios Krestenitis, Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas |
Pattern Recognit. | 4 |
| 2020 | Efficient adaptive inference for deep convolutional neural networks using hierarchical early exits
Nikolaos Passalis, Jenni Raitoharju, Anastasios Tefas, Moncef Gabbouj |
Pattern Recognit. | 4 |
| 2020 | Variance-preserving deep metric learning for content-based image retrievalabstractSupervised deep metric learning led to spectacular results for several Content-based Information Retrieval (CBIR) applications. The success of these approaches slowly led to the belief that image retrieval and classification are just slightly different variations of the same problem. However, recent evidence suggests that learning highly discriminative representation for a (limited) set of training classes removes valuable information from the representation, potentially harming both the in-domain, as well as the out-of-domain retrieval precision. In this paper, we propose a regularized discriminative deep metric learning method that aims to not only learn a representation that allows for discriminating between different classes, but it is also capable of encoding the latent generative factors separately for each class, overcoming this limitation. This allows for modeling the in-class variance and, as a result, maintaining the ability to represent both sub-classes of the in-domain data, as well as objects that belong to classes outside the training domain. The effectiveness of the proposed method, over existing supervised and unsupervised representation/metric learning approaches , is demonstrated under different in-domain and out-of-domain setups and three challenging image datasets. Nikolaos Passalis, Alexandros Iosifidis, Moncef Gabbouj, Anastasios Tefas |
Pattern Recognit. Lett. | 3 |
| 2020 | Temporal logistic neural Bag-of-Features for financial time series forecasting leveraging limit order book dataabstract• Logistic Neural Bag-of-Features are employed for financial time series analysis. • The proposed method can be efficiently used in deep learning architectures. • An adaptive scaling method is proposed to ensure the smooth flow of information. • A logistic kernel is used to estimate the feature vector densities. • The proposed method outperforms the competitive methods on a large-scale dataset. Time series forecasting is a crucial component of many important applications, ranging from forecasting the stock markets to energy load prediction. The high-dimensionality, velocity and variety of the data collected in many of these applications pose significant and unique challenges that must be carefully addressed for each of them. In this work, a novel Temporal Logistic Neural Bag-of-Features approach, that can be used to tackle these challenges, is proposed. The proposed method can be effectively combined with deep neural networks , leading to powerful deep learning models for time series analysis. However, combining existing BoF formulations with deep feature extractors pose significant challenges: the distribution of the input features is not stationary, tuning the hyper-parameters of the model can be especially difficult and the normalizations involved in the BoF model can cause significant instabilities during the training process. The proposed method is capable of overcoming these limitations by a employing a novel adaptive scaling mechanism and replacing the classical Gaussian-based density estimation involved in the regular BoF model with a logistic kernel. The effectiveness of the proposed approach is demonstrated using extensive experiments on a large-scale limit order book dataset that consists of more than 4 million limit orders. Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
Pattern Recognit. Lett. | 4 |
| 2020 | Human experts vs. machines in taxa recognition
Johanna Ärje, Jenni Raitoharju, Alexandros Iosifidis, Ville Tirronen, Kristian Meissner, Moncef Gabbouj, Serkan Kiranyaz, Salme Kärkkäinen |
Signal Process. Image Commun. | 6 |
| 2020 | DR-GAN: Automatic Radial Distortion Rectification Using Conditional GAN in Real-TimeabstractRadial distortion, which severely hinders object detection and semantic recognition, frequently exists in images captured using a wide-angle lens. Correction of this distortion of images is crucial in many computer vision applications. In this paper, we present distortion rectification generative adversarial network (DR-GAN), a conditional generative adversarial network (GAN) for automatic radial DR. To the best of our knowledge, this is the first end-to-end trainable adversarial framework for radial distortion rectification. The DR-GAN trained using the proposed low-to-high perceptual loss learns the mapping relation between different structural images rather than estimating multifarious distortion parameters, while also realizing label-free training and one-stage rectification. As a benefit of one-stage rectification, the proposed method is extremely fast with the completion of rectification in real time. This is approximately 22 times faster than the state-of-the-art methods. The experimental results show that the DR-GAN achieves an excellent performance in both quantitative measure (PSNR and SSIM) and visual qualitative appearance. Kang Liao, Chunyu Lin, Yao Zhao 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Distortion Rectification From Static to Dynamic: A Distortion Sequence Construction PerspectiveabstractDistortion rectification is a fundamental task in the field of computer vision and image processing. Nevertheless, previous methods have regarded distortion rectification as a static problem that learns a mapping function and corrects the distorted image to a unique state. However, this state is generally not the optimal solution, as it would result in an under-rectified or over-rectified structure. In this study, we revisit the classical distortion rectification task with a new perspective and redesign the algorithm, inspired by video processing techniques. Specifically, we regard distortion rectification as a dynamic problem that can be extended to a sequence of different distortion states: the input distorted image (t), under-rectified image (t+1), ideal-rectified image (t+2), and over-rectified image (t+3). We first estimate the residual distortion map (RDM) between the input distorted image and the coarse-rectified (t+1 or t+3) image. Here, RDM indicates the motion difference between two distorted images. Subsequently, the RDM is used to guide the refinement rectification process, aiming to convert the coarse-rectified state into the ideal-rectified state. In addition, the flexible implementation of the proposed refinement process with RDM to improve the rectification results of any method is appealing. The experimental results demonstrate that our method outperforms the state-of-the-art schemes by a significant margin, revealing approximately 40% improvement through quantitative evaluation. Kang Liao, Chunyu Lin, Yao Zhao 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Bag of Color Features for Color ConstancyabstractIn this paper, we propose a novel color constancy approach, called Bag of Color Features (BoCF), building upon Bag-of-Features pooling. The proposed method substantially reduces the number of parameters needed for illumination estimation. At the same time, the proposed method is consistent with the color constancy assumption stating that global spatial information is not relevant for illumination estimation and local information (edges, etc.) is sufficient. Furthermore, BoCF is consistent with color constancy statistical approaches and can be interpreted as a learning-based extension of many statistical approaches. To further improve the illumination estimation accuracy, we propose a novel attention mechanism for the BoCF model with two variants based on self-attention. BoCF approach and its variants achieve competitive, compared to the state of the art, results while requiring much fewer parameters on three benchmark datasets: ColorChecker RECommended, INTEL-TUT version 2, and NUS8. Firas Laakom, Nikolaos Passalis, Jenni Raitoharju, Jarno Nikkanen, Anastasios Tefas, Alexandros Iosifidis, Moncef Gabbouj |
IEEE Trans. Image Process. | 7 |
| 2020 | Patient-Specific Seizure Detection Using Nonlinear Dynamics and NullclinesabstractNonlinear dynamics has recently been extensively used to study epilepsy due to the complex nature of the neuronal systems. This study presents a novel method that characterizes the dynamic behavior of pediatric seizure events and introduces a systematic approach to locate the nullclines on the phase space when the governing differential equations are unknown. Nullclines represent the locus of points in the solution space where the components of the velocity vectors are zero. A simulation study over 5 benchmark nonlinear systems with well-known differential equations in three-dimensional exhibits the characterization efficiency and accuracy of the proposed approach that is solely based on the reconstructed solution trajectory. Due to their unique characteristics in the nonlinear dynamics of epilepsy, discriminative features can be extracted based on the nullclines concept. Using a limited training data (only 25% of each EEG record) in order to mimic the real-world clinical practice, the proposed approach achieves 91.15% average sensitivity and 95.16% average specificity over the benchmark CHB-MIT dataset. Together with an elegant computational efficiency, the proposed approach can, therefore, be an automatic and reliable solution for patient-specific seizure detection in long EEG recordings. Morteza Zabihi, Serkan Kiranyaz, Ville Jäntti, Tarmo Lipping, Moncef Gabbouj |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Deep Adaptive Input Normalization for Time Series ForecastingabstractDeep learning (DL) models can be used to tackle time series analysis tasks with great success. However, the performance of DL models can degenerate rapidly if the data are not appropriately normalized. This issue is even more apparent when DL is used for financial time series forecasting tasks, where the nonstationary and multimodal nature of the data pose significant challenges and severely affect the performance of DL models. In this brief, a simple, yet effective, neural layer that is capable of adaptively normalizing the input time series, while taking into account the distribution of the data, is proposed. The proposed layer is trained in an end-to-end fashion using backpropagation and leads to significant performance improvements compared to other evaluated normalization schemes. The proposed method differs from traditional normalization methods since it learns how to perform normalization for a given task instead of using a fixed normalization scheme. At the same time, it can be directly applied to any new time series without requiring retraining. The effectiveness of the proposed method is demonstrated using a large-scale limit order book data set, as well as a load forecasting data set. Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Heterogeneous Multilayer Generalized Operational PerceptronabstractThe traditional multilayer perceptron (MLP) using a McCulloch-Pitts neuron model is inherently limited to a set of neuronal activities, i.e., linear weighted sum followed by nonlinear thresholding step. Previously, generalized operational perceptron (GOP) was proposed to extend the conventional perceptron model by defining a diverse set of neuronal activities to imitate a generalized model of biological neurons. Together with GOP, a progressive operational perceptron (POP) algorithm was proposed to optimize a predefined template of multiple homogeneous layers in a layerwise manner. In this paper, we propose an efficient algorithm to learn a compact, fully heterogeneous multilayer network that allows each individual neuron, regardless of the layer, to have distinct characteristics. Based on the complexity of the problem, the proposed algorithm operates in a progressive manner on a neuronal level, searching for a compact topology, not only in terms of depth but also width, i.e., the number of neurons in each layer. The proposed algorithm is shown to outperform other related learning methods in extensive experiments on several classification problems. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | 1-D Convolutional Neural Networks for Signal Processing Applicationsabstract1D Convolutional Neural Networks (CNNs) have recently become the state-of-the-art technique for crucial signal processing applications such as patient-specific ECG classification, structural health monitoring, anomaly detection in power electronics circuitry and motor-fault detection. This is an expected outcome as there are numerous advantages of using an adaptive and compact 1D CNN instead of a conventional (2D) deep counterparts. First of all, compact 1D CNNs can be efficiently trained with a limited dataset of 1D signals while the 2D deep CNNs, besides requiring 1D to 2D data transformation, usually need datasets with massive size, e.g., in the "Big Data" scale in order to prevent the well-known "overfitting" problem. 1D CNNs can directly be applied to the raw signal (e.g., current, voltage, vibration, etc.) without requiring any pre- or post-processing such as feature extraction, selection, dimension reduction, denoising, etc. Furthermore, due to the simple and compact configuration of such adaptive 1D CNNs that perform only linear 1D convolutions (scalar multiplications and additions), a real-time and low-cost hardware implementation is feasible. This paper reviews the major signal processing applications of compact 1D CNNs with a brief theoretical background. We will present their state-of-the-art performances and conclude with focusing on some major properties. Serkan Kiranyaz, Turker Ince, Osama Abdeljaber, Onur Avci, Moncef Gabbouj |
ICASSP | 5 |
| 2019 | Deep Temporal Logistic Bag-of-features for Forecasting High Frequency Limit Order Book Time SeriesabstractForecasting time series has several applications in various domains. The vast amount of data that are available nowadays provide the opportunity to use powerful deep learning approaches, but at the same time pose significant challenges of high-dimensionality, velocity and variety. In this paper, a novel logistic formulation of the well-known Bag-of-Features model is proposed to tackle these challenges. The proposed method is combined with deep convolutional feature extractors and is capable of accurately modeling the temporal behavior of time series, forming powerful forecasting models that can be trained in an end-to-end fashion. The proposed method was extensively evaluated using a large-scale financial time series dataset, that consists of more than 4 million limit orders, outperforming other competitive methods. Nikolaos Passalis, Anastasios Tefas, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis |
ICASSP | 4 |
| 2019 | Adaptive Inference Using Hierarchical Convolutional Bag-of-Features for Low-Power Embedded PlatformsabstractUsing early exits provide a straightforward way to implement models that can adapt on-the-fly to the available computational resources. However, early exits in many cases suffer from significant limitations, which often prohibit their practical application, especially when placed on convolutional layers with narrow receptive fields. In this work, we propose a method capable of overcoming these limitations by a) using a Bag-of-Features (BoF)-based pooling approach, that allows for keeping more information regarding the distribution of the extracted feature vectors, while also maintaining more spatial information and b) employing a simple, yet effective, hierarchical approach for designing the exits, allowing for efficiently re-using the information that was already extracted by the previous layers. It is experimentally demonstrated that the proposed approach leads to significant performance improvements, allowing early exits to be a more practical tool that can be used in many real-world embedded applications. Nikolaos Passalis, Jenni Raitoharju, Anastasios Tefas, Moncef Gabbouj |
ICIP | 4 |
| 2019 | Class-Based Variational Representation Learning For Robust Image RetrievalabstractSupervised learning for Content-based Information Retrieval allows for obtaining discriminative representations that often excel within the training domain. However, recent evidence suggests that these representations can actually harm the retrieval precision for queries that do not belong to the domain of the training set compared to other, less discriminative representations. To avoid this behavior, we propose to learn discriminative representations which also encode the latent generative factors for each class. In this way, the proposed method is capable of maintaining (part of) the in-class variance, as well as being able to represent data that belong to classes that were not seen during the training by better learning the structure of the input space. The proposed method is evaluated under different in-domain and out-of-domain setups, significantly outperforming existing supervised and unsupervised representation learning approaches. Nikolaos Passalis, Anastasios Tefas, Alexandros Iosifidis, Moncef Gabbouj |
ICIP | 4 |
| 2019 | Knowledge Transfer for Face Verification Using Heterogeneous Generalized Operational PerceptronsabstractFace verification is a prominent biometric technique for identity authentication that has been used extensively in several security applications. In practice, face verification is often performed along with other visual surveillance tasks in the computing device. Thus, the ability to share the computation and reuse the information already extracted for other analysis tasks can greatly help reduce the computation load on the devices. In this study, we propose to utilize the knowledge transfer approach for the face verification problem by building a heterogeneous neural network architecture of Generalized Operational Perceptrons on top of the intermediate features extracted for object recognition purpose. Experimental results show that using our proposed approach, a face verification system can be incorporated into an existing visual analysis system with less additional memory and computational cost, compared to other similar approaches. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
ICIP | 3 |
| 2019 | PyGOP: A Python library for Generalized Operational Perceptron algorithms
Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
Knowl. Based Syst. | 3 |
| 2019 | Shared Coded Picture Technique for Tile-Based Viewport-Adaptive Streaming of Omnidirectional VideoabstractTile-based viewport-adaptive streaming methods have been used in delivering omnidirectional video for virtual reality applications. In these methods, the 360° video is encoded in multiple quality versions by using the motion constrained tile set (MCTS) technique. A set of high-quality and low-quality tiles, corresponding to viewport and non-viewport areas, respectively, are selected and transmitted to the user. However, these methods require frequent intra random access points to ensure seamless viewport switching capability, very high decoding complexity, or a multi-layer coding scheme. The frequent intra random access points include very high bitrate in viewport switching points. The high decoding complexity and multi-layer decoder requirements are not aligned with the omnidirectional media format (OMAF) standard. Such requirements make these methods sub-optimal or impractical for streaming the omnidirectional video. This paper studies the current tile-based solutions for delivering the omnidirectional content. Moreover, the OMAF-compliant shared coded picture (SCP)-based scheme is proposed in this paper for streaming the omnidirectional video. The core concept of the SCP-based method is to manipulate the switching point pictures in a way that the frequent intra-coded pictures are no longer required for the viewport switching operations between different quality versions of the content. The experiments illustrated that the SCP-based method outperforms the MCTS-based method on average by 11% to 14% in terms of streaming bitrate reduction with only 4% extra decoding complexity. Ramin Ghaznavi Youvalari, Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | Temporal Attention-Augmented Bilinear Network for Financial Time-Series Data AnalysisabstractFinancial time-series forecasting has long been a challenging problem because of the inherently noisy and stochastic nature of the market. In the high-frequency trading, forecasting for trading purposes is even a more challenging task, since an automated inference system is required to be both accurate and fast. In this paper, we propose a neural network layer architecture that incorporates the idea of bilinear projection as well as an attention mechanism that enables the layer to detect and focus on crucial temporal information. The resulting network is highly interpretable, given its ability to highlight the importance and contribution of each temporal instance, thus allowing further analysis on the time instances of interest. Our experiments in a large-scale limit order book data set show that a two-hidden-layer network utilizing our proposed layer outperforms by a large margin all existing state-of-the-art results coming from much deeper architectures while requiring far fewer computations. Dat Thanh Tran, Alexandros Iosifidis, Juho Kanniainen, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | 6K and 8K Effective Resolution with 4K HEVC Decoding Capability for 360 Video StreamingabstractThe recent Omnidirectional MediA Format (OMAF) standard, which specifies the delivery of 360° video content, supports only equirectangular projection (ERP) and cubemap projection and their region-wise packing with a limitation on video decoding capability to the maximum resolution of 4K (e.g., 4,096 × 2,048). Streaming of 4K ERP content allows only a limited viewport resolution, which is lower than the resolution of many current head-mounted displays (HMDs). Therefore, to take full advantage of high-resolution HMDs, delivery of 360° video content beyond 4K resolution needs to be enabled. In this regard, we propose two specific mixed-resolution packing schemes of 6K (e.g., 6,144 × 3,072) and 8K (e.g., 8,192 × 4,096) ERP content and their realization in tile-based streaming, while complying with the 4K decoding constraint and the High Efficiency Video Coding standard. The proposed packing schemes offer 6K and 8K effective resolution at the viewport. Using our proposed test methodology, experimental results indicate that the proposed layouts significantly decrease streaming bitrates when compared to mixed-quality viewport-adaptive streaming of 4K ERP. Our results further indicate that 8K-effective packing outperforms 6K-effective packing especially in high-quality videos. Alireza Zare, Maryam Homayouni, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2018 | Low-Energy Graph Fourier Basis Functions Span Salient ObjectsabstractThere is an emerging interest aiming at defining principles for signals on general graphs, which are analogous to the basic principles in traditional signal processing. One example is the Graph Fourier Transform which aims at decomposing a graph signal into its components based on a set of basis functions with corresponding graph frequencies. It has been observed that most of the important information of a graph signal is contained inside the low frequency band, which leads to several applications such as denoising, compression, etc. In this paper, we show that the low frequency basis functions span the salient regions in an image, which can also be considered as important regions. Motivated by this, we present a novel simple and unsupervised method to utilize a number of low-energy basis functions and show that it improves the performance of seven state-of-the-art salient object detection methods in five datasets under four different evaluation criteria, with only minor exceptions. Junaid Malik, Çaglar Aytekin, Moncef Gabbouj |
ICASSP | 3 |
| 2018 | Acceleration Approaches for Big Data AnalysisabstractThe massive size of data that needs to be processed by Machine Learning models nowadays sets new challenges related to their computational complexity and memory footprint. These challenges span all processing steps involved in the application of the related models, i.e., from the fundamental processing steps needed to evaluate distances of vectors, to the optimization of large-scale systems, e.g. for non-linear regression using kernels, or the speed up of deep learning models formed by billions of parameters. In order to address these challenges, new approximate solutions have been recently proposed based on matrix/tensor decompositions, randomization and quantization strategies. This paper provides a comprehensive review of the related methodologies and discusses their connections. Anton Muravev, Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis, Serkan Kiranyaz |
ICIP | 3 |
| 2018 | Weighted Linear Discriminant Analysis Based on Class Saliency InformationabstractIn this paper, we propose a new variant of Linear Discriminant Analysis to overcome underlying drawbacks of traditional LDA and other LDA variants targeting problems involving imbalanced classes. Traditional LDA sets assumptions related to Gaussian class distribution and neglects influence of outlier classes, that might hurt in performance. We exploit intuitions coming from a probabilistic interpretation of visual saliency estimation in order to define saliency of a class in multi-class setting. Such information is then used to redefine the between-class and within-class scatters in a more robust manner. Compared to traditional LDA and other weight-based LDA variants, the proposed method has shown certain improvements on facial image classification problems in publicly available datasets. Lei Xu 0036, Alexandros Iosifidis, Moncef Gabbouj |
ICIP | 3 |
| 2018 | Feature Dimensionality Reduction with Graph Embedding and Generalized Hamming DistanceabstractPrincipal component analysis (PCA) and linear discriminant analysis (LDA) are the most well-known methods to reduce the dimensionality of feature vectors. However, both methods face challenges when used on multilabel data - each data point may be associated to multiple labels. PCA does not take advantage of label information thus the performance is sacrificed. LDA can exploit class information for multiclass data, but cannot be directly applied to multilabel problems. In this paper, we propose a novel dimensionality reduction method for multilabel data. We first introduce the generalized Hamming distance that measures the distance of two data points in the label space. Then the proposed distance is used in the graph embedding framework for feature dimension reduction. We verified the proposed method using three multilabel benchmark datasets and one large image dataset. The results show that the proposed feature dimensionality reduction method consistently outperforms PCA and other competing methods. Honglei Zhang 0001, Moncef Gabbouj |
ICIP | 2 |
| 2018 | Subspace Support Vector Data DescriptionabstractThis paper proposes a novel method for solving one-class classification problems. The proposed approach, namely Subspace Support Vector Data Description, maps the data to a subspace that is optimized for one-class classification. In that feature space, the optimal hypersphere enclosing the target class is then determined. The method iteratively optimizes the data mapping along with data description in order to define a compact class representation in a low-dimensional feature space. We provide both linear and non-linear mappings for the proposed method. Experiments on 14 publicly available datasets indicate that the proposed Subspace Support Vector Data Description provides better performance compared to baselines and other recently proposed one-class classification methods. Fahad Sohrab, Jenni Raitoharju, Moncef Gabbouj, Alexandros Iosifidis |
ICPR | 3 |
| 2018 | Guest Editorial Special Issue on Multimedia Big Data in Internet of ThingsabstractMultimedia big data is one of the cornerstones of Internet of Things (IoT). IoT research subjects naturally connect with multimedia and big data. At the same time, multimedia data and application occupies a large proportion of the landscape of IoT. Moreover, there are many exciting new research directions in multimedia big data based IoT, such as directional sensor networks, video opportunistic transmission, physical object location, image or voice based physical object searching, and so on. Due to the dramatic development of IoT, we have bigger and bigger data sets, and we are stepping into the world of multimedia big data. Huadong Ma, Shui Yu 0001, Moncef Gabbouj, Peter Mueller |
IEEE Internet Things J. | 3 |
| 2018 | Benchmark database for fine-grained image classification of benthic macroinvertebrates
Jenni Raitoharju, Ekaterina Riabchenko, Iftikhar Ahmad 0001, Alexandros Iosifidis, Moncef Gabbouj, Serkan Kiranyaz, Ville Tirronen, Johanna Ärje, Salme Kärkkäinen, Kristian Meissner |
Image Vis. Comput. | 5 |
| 2018 | Feature synthesis for image classification and retrieval via one-against-all perceptrons
Jenni Raitoharju, Serkan Kiranyaz, Moncef Gabbouj |
Neural Comput. Appl. | 3 |
| 2018 | Improving efficiency in convolutional neural networks with multilinear filtersabstractThe excellent performance of deep neural networks has enabled us to solve several automatization problems, opening an era of autonomous devices. However, current deep net architectures are heavy with millions of parameters and require billions of floating point operations. Several works have been developed to compress a pre-trained deep network to reduce memory footprint and, possibly, computation. Instead of compressing a pre-trained network, in this work, we propose a generic neural network layer structure employing multilinear projection as the primary feature extractor. The proposed architecture requires several times less memory as compared to the traditional Convolutional Neural Networks (CNN), while inherits the similar design principles of a CNN. In addition, the proposed architecture is equipped with two computation schemes that enable computation reduction or scalability. Experimental results show the effectiveness of our compact projection that outperforms traditional CNN, while requiring far fewer parameters. Dat Thanh Tran, Alexandros Iosifidis, Moncef Gabbouj |
Neural Networks | 3 |
| 2018 | Probabilistic saliency estimation
Çaglar Aytekin, Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 3 |
| 2018 | Generalized Multi-View Embedding for Visual Recognition and Cross-Modal RetrievalabstractIn this paper, the problem of multi-view embedding from different visual cues and modalities is considered. We propose a unified solution for subspace learning methods using the Rayleigh quotient, which is extensible for multiple views, supervised learning, and nonlinear embeddings. Numerous methods including canonical correlation analysis, partial least square regression, and linear discriminant analysis are studied using specific intrinsic and penalty graphs within the same framework. Nonlinear extensions based on kernels and (deep) neural networks are derived, achieving better performance than the linear ones. Moreover, a novel multi-view modular discriminant analysis is proposed by taking the view difference into consideration. We demonstrate the effectiveness of the proposed multi-view embedding methods on visual object recognition and cross-modal image retrieval, and obtain superior results in both applications compared to related methods. Guanqun Cao, Alexandros Iosifidis, Ke Chen 0004, Moncef Gabbouj |
IEEE Trans. Cybern. | 4 |
| 2018 | A Data Set for Camera-Independent Color ConstancyabstractIn this paper, we provide a novel data set designed for Camera-independent color constancy research. Camera independence corresponds to the robustness of an algorithm's performance when it runs on images of the same scene taken by different cameras. Accordingly, the images in our database correspond to several laboratory and field scenes each of which is captured by three different cameras with minimal registration errors. The laboratory scenes are also captured under five different illuminations. The spectral responses of cameras and the spectral power distributions of the laboratory light sources are also provided, as they may prove beneficial for training future algorithms to achieve color constancy. For a fair evaluation of future methods, we provide guidelines for supervised methods with indicated training, validation, and testing partitions. Accordingly, we evaluate two recently proposed convolutional neural network-based color constancy algorithms as baselines for future research. As a side contribution, this data set also includes images taken by a mobile camera with color shading corrected and uncorrected results. This allows research on the effect of color shading as well. Çaglar Aytekin, Jarno Nikkanen, Moncef Gabbouj |
IEEE Trans. Image Process. | 3 |
| 2018 | Spatiotemporal Saliency Estimation by Spectral Foreground DetectionabstractWe present a novel approach for spatiotemporal saliency detection by optimizing a unified criterion of color contrast, motion contrast, appearance, and background cues. To this end, we first abstract the video by temporal superpixels. Second, we propose a novel graph structure exploiting the saliency cues to assign the edge weights. The salient segments are then extracted by applying a spectral foreground detection method, quantum cuts, on this graph. We evaluate our approach on several public datasets for video saliency and activity localization to demonstrate the favorable performance of the proposed video quantum cuts compared to the state of the art. Çaglar Aytekin, Horst Possegger, Thomas Mauthner, Serkan Kiranyaz, Horst Bischof, Moncef Gabbouj |
IEEE Trans. Multim. | 6 |
| 2017 | VisualLabel: An Integrated Multimedia Content Management and Access FrameworkabstractWith the rapid growth of image and video data as well as the fast spread of user-generated content in social media and cloud services, it has become increasingly difficult for users to have efficient access and effective management of their digital content. In this paper we present a novel integrated open source multimedia content management and access framework, called VisualLabel, that enables smart photo services based on automated visual content analysis, annotation, search and retrieval using state of the art analysis back ends for services such as Facebook and Flickr. This paper includes detailed descriptions of the high-level architecture used in the VisualLabel framework and proof-of-concept implementations of a front-end service, along with three analysis back ends and a web client, all of which demonstrate the basic functionality provided by the framework. Iftikhar Ahmad 0001, Petri Rantanen, Pekka Sillberg, Jorma Laaksonen, Thomas Forss, Aqdas Malik, Marko Nieminen, Rakshith Shetty, Satoru Ishikawa, Jarno Kallio, Jukka Saarinen, Moncef Gabbouj, Jari Soini |
EJC | 13 |
| 2017 | A k-nearest neighbor multilabel ranking algorithm with application to content-based image retrievalabstractMultilabel ranking is an important machine learning task with many applications, such as content-based image retrieval (CBIR). However, when the number of labels is large, traditional algorithms are either infeasible or show poor performance. In this paper, we propose a simple yet effective multilabel ranking algorithm that is based on k-nearest neighbor paradigm. The proposed algorithm ranks labels according to the probabilities of the label association using the neighboring samples around a query sample. Different from traditional approaches, we take only positive samples into consideration and determine the model parameters by directly optimizing ranking loss measures. We evaluated the proposed algorithm using four popular multilabel datasets. The proposed algorithm achieves equivalent or better performance than other instance-based learning algorithms. When applied to a CBIR system with a dataset of 1 million samples and over 190 thousand labels, which is much larger than any other multilabel datasets used earlier, the proposed algorithm clearly outperforms the competing algorithms. Honglei Zhang 0001, Serkan Kiranyaz, Moncef Gabbouj |
ICASSP | 3 |
| 2017 | Deep multi-resolution color constancyabstractIn this paper, a computational color constancy method is proposed via estimating the illuminant chromaticity in a scene by pooling from many local estimates. To this end, first, for each image in a dataset, we form an image pyramid consisting of several scales of the original image. Next, local patches of certain size are extracted from each scale in this image pyramid. Then, a convolutional neural network is trained to estimate the illuminant chromaticity per-patch. Finally, two more consecutive trainings are conducted, where the estimation is made per-image via taking the mean (1sttraining) and median (2ndtraining) of local estimates. The proposed method is shown to outperform the state-of-the-art in a widely used color constancy dataset. Çaglar Aytekin, Jarno Nikkanen, Moncef Gabbouj |
ICIP | 3 |
| 2017 | Category independent object proposals using quantum superpositionabstractObject proposals improve the efficiency of object detection by providing probable locations of objects in an image. Most of the state-of-the-art object proposal methods employ a supervised approach and learn object features from ground truth annotations. We present a novel unsupervised approach for generating object proposals that is based on the human visual system and quantum mechanical principles. Despite of being devoid of any learnt priors pertaining to objects in images, the proposed method is shown to yield competitive results with supervised approaches. Junaid Malik, Çaglar Aytekin, Moncef Gabbouj |
ICIP | 3 |
| 2017 | Class-specific kernel discriminant analysis based on Cholesky decompositionabstractIn this paper we describe a method for nonlinear class-specific discriminant learning that is based on Cholesky Decomposition. We show that the optimization problem solved in Class-Specific Kernel Discriminant Analysis is equivalent to that of Low-Rank Kernel Regression using training data independent target vectors. This connection allows us to devise a new Class-Specific Kernel Discriminant Analysis method that can be trained by exploiting fast linear system approaches, like the Cholesky decomposition. We verify our analysis in publicly available verification problems designed for human action recognition. Alexandros Iosifidis, Moncef Gabbouj |
IJCNN | 2 |
| 2017 | Generalized model of biological neural networks: Progressive operational perceptronsabstractTraditional Artificial Neural Networks (ANNs) such as Multi-Layer Perceptrons (MLPs) and Radial Basis Functions (RBFs) were designed to simulate biological neural networks; however, they are based only loosely on biology and only provide a crude model. This in turn yields well-known limitations and drawbacks on the performance and robustness. In this paper we shall address them by introducing a novel feed-forward ANN model, Generalized Operational Perceptrons (GOPs) that consist of neurons with distinct (non-)linear operators to achieve a generalized model of the biological neurons and ultimately a superior diversity. We modified the conventional back-propagation (BP) to train GOPs and furthermore, proposed Progressive Operational Perceptrons (POPs) to achieve self-organized and depth-adaptive GOPs according to the learning problem. The most crucial property of the POPs is their ability to simultaneously search for the optimal operator set and train each layer individually. The final POP is, therefore, formed layer by layer and this ability enables POPs with minimal network depth to attack the most challenging learning problems that cannot be learned by conventional ANNs even with a deeper and significantly complex configuration. Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
IJCNN | 4 |
| 2017 | Comparison of HEVC coding schemes for tile-based viewport-adaptive streaming of omnidirectional videoabstractVirtual reality applications make use of 360-degree panoramic or omnidirectional video with high resolution and high frame rate in order to create the immersive experience to the user. The user views only a portion of the captured 360-degree scene at each time instant, hence streaming the whole omnidirectional video in highest quality is not efficient. In order to alleviate the problem of bandwidth wastage, viewport-adaptive encoding and streaming schemes have been proposed. In these schemes, part of the captured scene that is within the viewer's field of view is delivered at highest quality while the rest of the scene in a lower quality. In this work, three tile-based viewport-adaptive methods using motion-constrained tile sets (MCTS), region-of-interest scalability and simulcast approach have been studied for streaming omnidirectional content. In the performed experiments with various tiling arrangements, MCTS-based scheme required highest bitrate compared to other methods. The scalable coding scheme provided the highest performance in terms of streaming bitrate saving on average up to 53% and 35% compared to streaming the whole omnidirectional video and MCTS-based method, respectively. Ramin Ghaznavi Youvalari, Alireza Zare, Huameng Fang, Alireza Aminlou, Qingpeng Xie, Miska M. Hannuksela, Moncef Gabbouj |
MMSP | 7 |
| 2017 | The effect of automated taxa identification errors on biological indices
Johanna Ärje, Salme Kärkkäinen, Kristian Meissner, Alexandros Iosifidis, Turker Ince, Moncef Gabbouj, Serkan Kiranyaz |
Expert Syst. Appl. | 6 |
| 2017 | On the comparison of random and Hebbian weights for the training of single-hidden layer feedforward neural networks
Kaveh Samiee, Alexandros Iosifidis, Moncef Gabbouj |
Expert Syst. Appl. | 3 |
| 2017 | Progressive Operational Perceptrons
Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neurocomputing | 4 |
| 2017 | CNN-based edge filtering for object proposals
Muhammad-Adeel Waris, Alexandros Iosifidis, Moncef Gabbouj |
Neurocomputing | 3 |
| 2017 | Epileptic seizure detection in long-term EEG records using sparse rational decomposition and local Gabor binary patterns feature extraction
Kaveh Samiee, Péter Kovács 0001, Moncef Gabbouj |
Knowl. Based Syst. | 3 |
| 2017 | Extended quantum cuts for unsupervised salient object extraction
Çaglar Aytekin, Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
Multim. Tools Appl. | 4 |
| 2017 | An effective CU size decision method for quality scalability in SHVC
Xiaoni Li, Mianshu Chen, Zhaowei Qu, Jimin Xiao, Moncef Gabbouj |
Multim. Tools Appl. | 5 |
| 2017 | Modeling the timing of cuts in automatic editing of concert videos
Mikko Roininen, Jussi Leppänen, Antti J. Eronen, Igor D. D. Curcio, Moncef Gabbouj |
Multim. Tools Appl. | 5 |
| 2017 | Urban 3D segmentation and modelling from street view images and LiDAR point cloudsabstract3D urban maps with semantic labels and metric information are not only essential for the next generation robots such autonomous vehicles and city drones, but also help to visualize and augment local environment in mobile user applications. The machine vision challenge is to generate accurate urban maps from existing data with minimal manual annotation. In this work, we propose a novel methodology that takes GPS registered LiDAR (Light Detection And Ranging) point clouds and street view images as inputs and creates semantic labels for the 3D points clouds using a hybrid of rule-based parsing and learning-based labelling that combine point cloud and photometric features. The rule-based parsing boosts segmentation of simple and large structures such as street surfaces and building facades that span almost 75% of the point cloud data. For more complex structures, such as cars, trees and pedestrians, we adopt boosted decision trees that exploit both structure (LiDAR) and photometric (street view) features. We provide qualitative examples of our methodology in 3D visualization where we construct parametric graphical models from labelled data and in 2D image segmentation where 3D labels are back projected to the street view images. In quantitative evaluation we report classification accuracy and computing times and compare results to competing methods with three popular databases: NAVTEQ True, Paris-Rue-Madame and TLS (terrestrial laser scanned) Velodyne. Pouria Babahajiani, Lixin Fan, Joni-Kristian Kämäräinen, Moncef Gabbouj |
Mach. Vis. Appl. | 4 |
| 2017 | Learning graph affinities for spectral graph-based salient object detection
Çaglar Aytekin, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
Pattern Recognit. | 4 |
| 2017 | Multilinear class-specific discriminant analysisabstractThere has been a great effort to transfer linear discriminant techniques that operate on vector data to high-order data, generally referred to as Multilinear Discriminant Analysis (MDA) techniques. Many existing works focus on maximizing the inter-class variances to intra-class variances defined on tensor data representations. However, there has not been any attempt to employ class-specific discrimination criteria for the tensor data. In this paper, we propose a multilinear subspace learning technique suitable for applications requiring class-specific tensor models. The method maximizes the discrimination of each individual class in the feature space while retains the spatial structure of the input. We evaluate the efficiency of the proposed method on two problems, i.e. facial image analysis and stock price prediction based on limit order book data. Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis |
Pattern Recognit. Lett. | 2 |
| 2017 | Big Media Data Analysis
Alexandros Iosifidis, Anastasios Tefas, Ioannis Pitas, Moncef Gabbouj |
Signal Process. Image Commun. | 4 |
| 2017 | Multi-View Nonparametric Discriminant Analysis for Image Retrieval and RecognitionabstractA novel multi-view nonparametric discriminant analysis method is proposed for the application of cross-modal image retrieval and zero-shot recognition. We exploit the class boundary structure and discrepancy information of the available views in order to formulate an optimization criterion, which is automatically adjusted to the multi-view class structures. The proposed method allows for multiple projection directions, by relaxing the Gaussian distribution assumption of related methods. The experiments demonstrate that the proposed method can achieve superior results comparing to several existing methods. Guanqun Cao, Alexandros Iosifidis, Moncef Gabbouj |
IEEE Signal Process. Lett. | 3 |
| 2017 | Row-Interleaved Sampling for Depth-Enhanced 3D Video Coding for Polarized DisplaysabstractPassive stereoscopic displays create the illusion of three dimensions by employing orthogonal polarizing filters and projecting two images onto the same screen. In this article, a coding scheme targeting depth-enhanced stereoscopic video coding for polarized displays is introduced. We propose to use asymmetric row-interleaved sampling for texture and depth views prior to encoding. The performance of the proposed scheme is compared with several other schemes, and the objective results confirm the superior performance of the proposed method. Furthermore, subjective evaluation proves that no quality degradation is introduced by the proposed coding scheme compared to the reference method. Maryam Homayouni, Payman Aflaki, Miska M. Hannuksela, Moncef Gabbouj |
ACM Trans. Appl. Percept. | 4 |
| 2017 | A New Block-Based Method for HEVC Intra CodingabstractThis paper presents a new block-based method for the High Efficiency Video Coding (HEVC) intra coding. First, we have found through analysis and test that the prediction errors on some pixels in each prediction block (PB) that are neighboring to the reference pixels would be no bigger than the corresponding coding errors. Based on this observation, the pixels in each PB are divided into two parts: half pixels are coded via a novel padding technique together with a constrained quantization algorithm (leading to around 3 dB gain under the same bit rate), whereas the other half are reconstructed by linear interpolations along a prediction direction by utilizing the neighboring reference pixels and the first half coded pixels. In the final implementation, a competition mechanism is employed between this new method and the original HEVC intra coding in order to choose the best mode for each PB. Experimental results show that about 2% BD-rate reduction has been achieved both for luma and chroma with respect to the original HEVC intra coding, whereas the encoder complexity increases by 130%, but the decoding time remains nearly unchanged. Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | A Hybrid Approach for Near-Range Video StabilizationabstractNear-range videos contain objects that are close to the camera. These videos often contain discontinuous depth variation (DDV), which is the main challenge to the existing video stabilization methods. Traditionally, 2D methods are robust to various camera motions (e.g., quick rotation and zooming) under scenes with continuous depth variation (CDV). However, in the presence of DDV, they often generate wobbled results due to the limited ability of their 2D motion models. Alternatively, 3D methods are more robust in handling near-range videos. We show that, by compensating rotational motions and ignoring translational motions, near-range videos can be successfully stabilized by 3D methods without sacrificing the stability too much. However, it is time-consuming to reconstruct the 3D structures for the entire video and sometimes even impossible due to rapid camera motions. In this paper, we combine the advantages of 2D and 3D methods, yielding a hybrid approach that is robust to various camera motions and can handle the near-range scenarios well. To this end, we automatically partition the input video into CDV and DDV segments. Then, the 2D and 3D approaches are adopted for CDV and DDV clips, respectively. Finally, these segments are stitched seamlessly via a constrained optimization. We validate our method on a large variety of consumer videos. Shuaicheng Liu, Binhan Xu, Chuang Deng, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2017 | Class-Specific Kernel Discriminant Analysis Revisited: Further Analysis and ExtensionsabstractIn this paper, we revisit class-specific kernel discriminant analysis (KDA) formulation, which has been applied in various problems, such as human face verification and human action recognition. We show that the original optimization problem solved for the determination of class-specific discriminant projections is equivalent to a low-rank kernel regression (LRKR) problem using training data-independent target vectors. In addition, we show that the regularized version of class-specific KDA is equivalent to a regularized LRKR problem, exploiting the same targets. This analysis allows us to devise a novel fast solution. Furthermore, we devise novel incremental, approximate and deep (hierarchical) variants. The proposed methods are tested in human facial image and action video verification problems, where their effectiveness and efficiency is shown. Alexandros Iosifidis, Moncef Gabbouj |
IEEE Trans. Cybern. | 2 |
| 2016 | Supervised subspace learning based on deep randomized networksabstractIn this paper, we propose a supervised subspace learning method that exploits the rich representation power of deep feedforward networks. In order to derive a fast, yet efficient, learning scheme we employ deep randomized neural networks that have been recently shown to provide good compromise between training speed and performance. For optimally determining the learnt subspace, we formulate a regression problem where we employ target vectors designed to encode both the labeling information available for the training data and geometric properties of the training data, when represented in the feature space determined by the network's last hidden layer outputs. We experimentally show that the proposed approach is able to outperform deep randomized neural networks trained by using the standard network target vectors. Alexandros Iosifidis, Moncef Gabbouj |
ICASSP | 2 |
| 2016 | Fisheye video coding using elastic motion compensated reference framesabstractFisheye cameras have become extremely popular in applications where the goal is to capture large fields of view with only one camera. However, the wide-angle fisheye imagery has special characteristics that may not be very well suited for modern video codecs that employ block-based translational motion model. This model fails to describe complex deformable motion which is often present in fisheye videos. In this paper, we advocate for the usage of elastic motion model in compensating such a complex motion. The presented design enables the re-use of existing codecs, such as HEVC, without modifications in low-level coding tools. Experimental results show that a savings in bit rate of up to 6.54% is achievable over standalone HEVC if the elastic motion compensated prediction is used as an additional reference frame. Ashek Ahmmed, Miska M. Hannuksela, Moncef Gabbouj |
ICIP | 3 |
| 2016 | Combining multi-class maximum margin classification with linear discriminant analysis for human action recognitionabstractIn this paper, a new multi-class classification method is proposed and evaluated in the problem of human action recognition in unconstrained environments. The proposed method exploits both the maximum margin property of multi-class Support Vector Machines and Linear Discriminant Analysis-based discrimination. Experiments indicate that by exploiting such discriminant information in a multi-class maximum margin framework, classification performance can be enhanced, leading to state-of-the-art performance in human action recognition. Alexandros Iosifidis, Moncef Gabbouj |
ICIP | 2 |
| 2016 | Face segmentation in thumbnail images by data-adaptive convolutional segmentation networksabstractIn this study we address the problem of face segmentation in thumbnail images. While there have been several approaches for face detection, none performs detection in such low resolution and segmentation with pixel accuracy. In this paper, we propose convolutional segmentation networks (CSNs) that can be trained to learn segmentation of human faces. Unlike the deep classifiers such as Convolutional Neural Network (CNNs), CSNs have the unique design solely for segmentation with minimal complexity. Furthermore, we propose a self-data organization (SDO) in order to create “expert” CSNs each of which is specialized over a set of images with certain face characteristics. SDO is integrated with CSN training in an interleaved manner and it is the key for the learning with simple and compact networks rather than the deep ones. This is especially a desired property for the limited face datasets with challenging face variations and complexities. Evaluations on the benchmark dataset show that CSNs can achieve an elegant segmentation accuracy despite the limited training data size, thumbnail resolution and highly complex face modalities. Serkan Kiranyaz, Muhammad-Adeel Waris, Iftikhar Ahmad 0001, Ridha Hamila, Moncef Gabbouj |
ICIP | 5 |
| 2016 | Salient object segmentation based on linearly combined affinity graphsabstractIn this paper, we propose a graph affinity learning method for a recently proposed graph-based salient object detection method, namely Extended Quantum Cuts (EQCut). We exploit the fact that the output of EQCut is differentiable with respect to graph affinities, in order to optimize linear combination coefficients and parameters of several differentiable affinity functions by applying error backpropagation. We show that the learnt linear combination of affinities improves the performance over the baseline method and achieves comparable (or even better) performance when compared to the state-of-the-art salient object segmentation methods. Çaglar Aytekin, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 4 |
| 2016 | Joint K-Means quantization for Approximate Nearest Neighbor SearchabstractRecently, Approximate Nearest Neighbor (ANN) Search has become a very popular approach for similarity search on large-scale datasets. In this paper, we propose a novel vector quantization method for ANN, which introduces a joint multi-layer K-Means clustering solution for determination of the codebooks. The performance of the proposed method is improved further by a joint encoding scheme. Experimental results verify the success of the proposed algorithm as it outperforms the state-of-the-art methods. Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 3 |
| 2016 | Learned vs. engineered features for fine-grained classification of aquatic macroinvertebratesabstractAquatic macroinvertebrate biomonitoring is an efficient way of assessment of slow and subtle anthropogenic changes and their effect on water quality. It is imperative to have reliable identification and counts of the various taxa occurring in samples as these form the basis for the quality indices used to infer the ecological status of the aquatic ecosystem. In this paper, we try to close the gap between human taxa identification accuracy (typically 90-95% on 30-40 classes of macroinvertebrates) and results of automatic fine-grained classification by introducing a novel technique based on Convolutional Neural Networks (CNN). CNN learns optimal features for macroinvertebrate classification and achieves near human accuracy when tested on 29 macroinvertebrate classes. Moreover, we perform comparative evaluation of the learned features against the hand-crafted features, which have been commonly used in classical approaches, and confirm superiority of the learned deep features over the engineered ones. Ekaterina Riabchenko, Kristian Meissner, Iftikhar Ahmad 0001, Alexandros Iosifidis, Ville Tirronen, Moncef Gabbouj, Serkan Kiranyaz |
ICPR | 6 |
| 2016 | Object proposals using CNN-based edge filteringabstractWith the success of deep learning in the last few years, the object detection community shifted from processing on exhaustive sliding windows to smaller set of object proposals using more powerful and deep visual representations. Object proposals increase the accuracy and speed up detection process by reducing the search space. In this paper we propose a novel idea of filtering irrelevant edges using semantic image filtering and true objectness learnt within convolutional layers of CNN. Our approach localizes well proposals by producing highly accurate bounding boxes and reduces the number of proposals. The greatest benefit of our approach is that it can be integrated into any existing method exploiting edge-based objectness to achieve consistently high recall across various intersection over union thresholds. Unlike other supervised methods, our approach does not require bounding box annotations for training. Experiments on PASCAL VOC 2007 dataset demonstrate that our approach improves the state-of-the-art model with a significant margin. Muhammad-Adeel Waris, Alexandros Iosifidis, Moncef Gabbouj |
ICPR | 3 |
| 2016 | An Optimized k-NN Approach for Classification on Imbalanced Datasets with Missing Data
Ezgi C. Ozan, Ekaterina Riabchenko, Serkan Kiranyaz, Moncef Gabbouj |
IDA | 4 |
| 2016 | Standard-Compliant Multiview Video Coding and Streaming for Virtual Reality ApplicationsabstractVirtual reality (VR) systems employ multiview cameras or camera rigs to capture a scene from the entire 360-degree perspective. Due to computational or latency constraints, it might not be possible to stitch multiview videos into a single video sequence prior to encoding. In this paper we investigate the coding and streaming of multiview VR video content. We present a standard-compliant method where we first divide the camera views into two types: Primary views represent a subset of camera views with lower resolution and non-overlapping (minimally overlapping) content which cover the entire 360-degree field-of-view to guarantee immediate monoscopic viewing during very rapid head movements. Auxiliary views consist of remaining camera views with higher resolution which produce overlapping content with the primary views and are additionally used for stereoscopic viewing. Based on this categorization, we propose a coding arrangement in which, the primary views are independently coded in the base layer and the additional auxiliary views are coded as an enhancement layer, using inter-layer prediction from primary views. The proposed system not only meets the low latency requirements of VR systems, but also conforms to the existing multilayer extensions of the High Efficiency Video Coding standard. Simulation results show that the coding and streaming performance of the proposed scheme is significantly improved compared to earlier methods. Kashyap Kammachi Sreedhar, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
ISM | 4 |
| 2016 | Viewport-Adaptive Encoding and Streaming of 360-Degree Video for Virtual Reality ApplicationsabstractVirtual reality applications use 360-degree videos and head mount displays (HMDs) with stereoscopic capabilities to provide full immersion experience. In these applications it is also common to use 4K resolution or higher per view for 360-degree videos. Consequently, this leads to technical challenges in handling the bandwidth requirements while keeping the system latency to the minimal. When the content is viewed with a HMD, a subset of the entire 360-degree video is displayed at a single point of time. To improve the resolution and picture quality of the displayed content, viewport based coding is desirable. In this regard, we investigated various viewport dependent projection schemes including the existing variants of Pyramidal projection. In this regard we propose the multi-resolution versions of Equirectangular and Cubemap projections. Additionally, we developed a methodology for comparing the rate-distortion performance of these projections. Based on the simulation results, it was observed that multi-resolution projections of Equirectangle and Cubemap outperform other projection schemes, significantly. Kashyap Kammachi Sreedhar, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
ISM | 4 |
| 2016 | Efficient Coding of 360-Degree Pseudo-Cylindrical Panoramic Video for Virtual Reality ApplicationsabstractPseudo-cylindrical panoramas represent the data distribution of spherical coordinates closely in two-dimensional domain due to the equidistant sampling of 360-degree scene. Therefore, unlike the cylindrical projections, they do not suffer from the over stretching in the polar areas. However, due to the non-rectangular format in effective picture area and sharp edges at its borders, the compression performance is inefficient. In this paper, we propose two methods which improve the compression performance of both intra-frame and inter-frame coding of pseudo-cylindrical panoramic content and meanwhile reduce the coding artifacts. In the intra-frame coding method, border edges are smoothed by modifying the content of the image in the non-effective picture area, which are cropped at the receiver side. In the inter-frame coding method, gaining the benefit of 360-degree property of the content, non-effective picture area of reference frames at border is filled with the content of the effective picture area from the opposite border to enhance the performance of motion compensation. Ramin Ghaznavi Youvalari, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
ISM | 4 |
| 2016 | HEVC-compliant Tile-based Streaming of Panoramic Video for Virtual Reality ApplicationsabstractDelivering wide-angle and high-resolution spherical panoramic video content entails a high streaming bitrate. This imposes challenges when panorama clips are consumed in virtual reality (VR) head-mounted displays (HMD). The reason is that the HMDs typically require high spatial and temporal fidelity contents and strict low-latency in order to guarantee the user's sense of presence while using them. In order to alleviate the problem, we propose to store two versions of the same video content at different resolutions, each divided into multiple tiles using the High Efficiency Video Coding (HEVC) standard. According to the user's present viewport, a set of tiles is transmitted in the highest captured resolution, while the remaining parts are transmitted from the low-resolution version of the same content. In order to enable randomly choosing different combinations, the tile sets are encoded to be independently decodable. We further study the trade-off in the choice of tiling scheme and its impact on compression and streaming bitrate performances. The results indicate streaming bitrate saving from 30% to 40%, depending on the selected tiling scheme, when compared to streaming the entire video content. Alireza Zare, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
ACM Multimedia | 4 |
| 2016 | HEVC-compliant viewport-adaptive streaming of stereoscopic panoramic videoabstractVirtual reality (VR) provides unprecedented immersive experience using high-resolution spherical stereoscopic panoramic video. Such an experience is achieved by using head-mounted display (HMD) which has very strict latency bounds in order to respond promptly to user movements. Conventional streaming of VR video requires large bandwidth because the entire captured panorama is transmitted. However, only a limited field-of-view (FOV) is displayed by an HMD, resulting in wastage of bandwidth. To alleviate the problem, this paper proposes a High Efficiency Video Coding (HEVC) compliant approach for efficient coding and streaming of stereoscopic VR content. The proposed method is based on partitioning video pictures into tiles, where only the required tiles corresponding to the primary viewport are transmitted in high resolution, while the remaining parts are transmitted in low resolution. Furthermore, this method enables coding stereoscopic video contents using a conventional HEVC codec, while still achieving significant compression gain by means of adopting inter-view prediction only in intra random access point (IRAP) pictures. Using this method, the predicted view can be decoded independently of the main view, hence allowing simultaneous decoding instances. Experimental results demonstrate that the proposed approach is able to substantially improve compression efficiency and streaming bitrate performance. Alireza Zare, Kashyap Kammachi Sreedhar, Vinod Kumar Malamal Vadakital, Alireza Aminlou, Miska M. Hannuksela, Moncef Gabbouj |
PCS | 6 |
| 2016 | Automated tree detection and density calculation using unmanned aerial vehiclesabstractGrowth monitoring in tree plantation farms plays a crucial role in taking necessary actions from early stages to improve harvesting efficiency. Unmanned aerial vehicles (UAVs) provide a fast and efficient way to acquire data from large farms with difficult access. We propose a novel end-to-end system composed of multiple stages to identify areas where seeds have not grown as expected in young tree farms. The system acquires data from UAV flights to generate georeferenced orthophoto images. Next, local binary patterns, distance transform, and watershed segmentation methods are applied on images to detect trees. Finally, tree density distributions are calculated in the granularity of approximately 90 m2tiles. We assessed the system performance by comparing detection results against a ground truth set of over 50000 trees from 16 orthophoto images for two tree species. The median deviation was 2 trees per tile for both species, where the median tree count was 14 and 9 per tile, respectively. Olcay Guldogan, J. Rotola-Pukkila, Uvaraj Balasundaram, Thanh-Hai Le, Kamal Mannar, Taufan Mega Chrisna, Moncef Gabbouj |
VCIP | 7 |
| 2016 | Hierarchical class-specific kernel discriminant analysis for face verificationabstractIn this paper, a new method for nonlinear class-specific data projection is proposed for verification problems. We apply a hierarchical process formed by multiple nonlinear class-specific data projection layers in order to determine data representations in multiple subspaces enhancing class discrimination. We evaluate the proposed method on four publicly available facial image datasets and compare its performance with related methods and show its effectiveness. Alexandros Iosifidis, Moncef Gabbouj |
VCIP | 2 |
| 2016 | Multi-class Support Vector Machine classifiers using intrinsic and penalty graphs
Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 2 |
| 2016 | Nyström-based approximate kernel subspace learning
Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. | 2 |
| 2016 | Learning to rank salient segments extracted by multispectral Quantum Cuts
Çaglar Aytekin, Serkan Kiranyaz, Moncef Gabbouj |
Pattern Recognit. Lett. | 3 |
| 2016 | New R-D Optimization Criterion for Fast Mode Decision Algorithms in Video Coding and TransratingabstractMode decision has a significant effect on the quality and complexity of video coding. It is even more challenging when generating multiple bitstreams with different bitrates (BRs) in, for example, dynamic adaptive streaming for HTTP or transrating systems. Full search and simplified fast search mode decision (MD) methods either suffer from a high computational complexity or have a negative impact on quality. Furthermore, mode selection in conventional approaches strongly depends on the quantization parameter (QP). Hence, modes that have been selected for high BR compression may not be suitable for low BR when transrating a bitstream. In this paper, we propose a rate-distortion (R-D)-optimized criterion for fast MD algorithms. The proposed cost function, when adopted in different fast MD algorithms, not only improves the R-D performance by up to 6.6% in terms of Bjøntegaard delta rate, but also reduces the execution time of the encoder by up to 6.8%. We also show that modes selected by the proposed criterion are less sensitive to changes in BR or QP. As a result, the same modes in an encoded bitstream may be used even after transrating using requantization, resulting in a significant R-D performance improvement of up to 33.3%. Alireza Aminlou, Mahmoud Reza Hashemi, Moncef Gabbouj, Bing Zeng 0001, Omid Fatemi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Scaling Up Class-Specific Kernel Discriminant Analysis for Large-Scale Face VerificationabstractIn this paper, a novel approximate solution of the criterion used in non-linear class-specific discriminant subspace learning is proposed. We build on the class-specific kernel spectral regression method, which is a two-step process formed by an eigenanalysis step and a kernel regression step. Based on the structure of the intra-class and out-of-class scatter matrices, we provide a fast solution for the first step. For the second step, we propose the use of approximate kernel space definitions. We analytically show that the adoption of randomized and class-specific kernels has the effect of regularization and Nyström-based approximation, respectively. We evaluate the proposed approach in face verification problems and compare it with the existing approaches. Experimental results show the effectiveness and efficiency of the proposed approximate class-specific kernel spectral regression method, since it can provide satisfactory performance and scale well with the size of the data. Alexandros Iosifidis, Moncef Gabbouj |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2016 | Joint Video Stitching and Stabilization From Moving CamerasabstractIn this paper, we extend image stitching to video stitching for videos that are captured for the same scene simultaneously by multiple moving cameras. In practice, videos captured under this circumstance often appear shaky. Directly applying image stitching methods for shaking videos often suffers from strong spatial and temporal artifacts. To solve this problem, we propose a unified framework in which video stitching and stabilization are performed jointly. Specifically, our system takes several overlapping videos as inputs. We estimate both inter motions (between different videos) and intra motions (between neighboring frames within a video). Then, we solve an optimal virtual 2D camera path from all original paths. An enlarged field of view along the virtual path is finally obtained by a space-temporal optimization that takes both inter and intra motions into consideration. Two important components of this optimization are that: 1) a grid-based tracking method is designed for an improved robustness, which produces features that are distributed evenly within and across multiple views and 2) a mesh-based motion model is adopted for the handling of the scene parallax. Some experimental results are provided to demonstrate the effectiveness of our approach on various consumer-level videos and a Plugin, named "Video Stitcher" is developed at Adobe After Effects CC2015 to show the processed videos. Heng Guo 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Image Process. | 6 |
| 2016 | K-Subspaces Quantization for Approximate Nearest Neighbor SearchabstractApproximate Nearest Neighbor (ANN) search has become a popular approach for performing fast and efficient retrieval on very large-scale datasets in recent years, as the size and dimension of data grow continuously. In this paper, we propose a novel vector quantization method for ANN search which enables faster and more accurate retrieval on publicly available datasets. We define vector quantization as a multiple affine subspace learning problem and explore the quantization centroids on multiple affine subspaces. We propose an iterative approach to minimize the quantization error in order to create a novel quantization scheme, which outperforms the state-of-the-art algorithms. The computational cost of our method is also comparable to that of the competing methods. Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Competitive Quantization for Approximate Nearest Neighbor SearchabstractIn this study, we propose a novel vector quantization algorithm for Approximate Nearest Neighbor (ANN) search, based on a joint competitive learning strategy and hence called as competitive quantization (CompQ). CompQ is a hierarchical algorithm, which iteratively minimizes the quantization error by jointly optimizing the codebooks in each layer, using a gradient decent approach. An extensive set of experimental results and comparative evaluations show that CompQ outperforms the-state-of-the-art while retaining a comparable computational complexity. Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Image Interpolation Based on Non-local Geometric Similarities and Directional GradientsabstractImage interpolation offers an efficient way to compose a high-resolution (HR) image from the observed low-resolution (LR) image. Advanced interpolation techniques design the interpolation weighting coefficients by solving a minimum mean-square-error (MMSE) problem in which the local geometric similarity is often considered. However, using local geometric similarities cannot usually make the MMSE-based interpolation as reliable as expected. To solve this problem, we propose a robust interpolation scheme by using the nonlocal geometric similarities to construct the HR image. In our proposed method, the MMSE-based interpolation weighting coefficients are generated by solving a regularized least squares problem that is built upon a number of dual-reference patches drawn from the given LR image and regularized by the directional gradients of these patches. Experimental results demonstrate that our proposed method offers a remarkable quality improvement as compared to some state-of-the-art methods, both objectively and subjectively. Shuyuan Zhu, Bing Zeng 0001, Liaoyuan Zeng, Moncef Gabbouj |
IEEE Trans. Multim. | 4 |
| 2016 | Training Radial Basis Function Neural Networks for Classification via Class-Specific ClusteringabstractIn training radial basis function neural networks (RBFNNs), the locations of Gaussian neurons are commonly determined by clustering. Training inputs can be clustered on a fully unsupervised manner (input clustering), or some supervision can be introduced, for example, by concatenating the input vectors with weighted output vectors (input-output clustering). In this paper, we propose to apply clustering separately for each class (class-specific clustering). The idea has been used in some previous works, but without evaluating the benefits of the approach. We compare the class-specific, input, and input-output clustering approaches in terms of classification performance and computational efficiency when training RBFNNs. To accomplish this objective, we apply three different clustering algorithms and conduct experiments on 25 benchmark data sets. We show that the class-specific approach significantly reduces the overall complexity of the clustering, and our experimental results demonstrate that it can also lead to a significant gain in the classification performance, especially for the networks with a relatively few Gaussian neurons. Among other applied clustering algorithms, we combine, for the first time, a dynamic evolutionary optimization method, multidimensional particle swarm optimization, and the class-specific clustering to optimize the number of cluster centroids and their locations. Jenni Raitoharju, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Visual saliency by extended quantum cutsabstractIn this study, we propose an unsupervised, state-of-the-art saliency map generation algorithm which is based on a recently proposed link between quantum mechanics and spectral graph clustering, Quantum Cuts. The proposed algorithm forms a graph among superpixels extracted from an image and optimizes a criterion related to the image boundary, local contrast and area information. Furthermore, the effects of the graph connectivity, superpixel shape irregularity, superpixel size and how to determine the affinity between superpixels are analyzed in detail. Furthermore, we introduce a novel approach to propose several saliency maps. Resulting saliency maps consistently achieves a state-of-the-art performance in a large number of publicly available benchmark datasets in this domain, containing around 18k images in total. Çaglar Aytekin, Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
ICIP | 4 |
| 2015 | Upsampled-view distortion optimization for mixed resolution 3D Video CodingabstractThe MVC+D extension of the Advanced Video Coding (H.264/AVC) standard enables multiview-and-depth 3D video coding but specifies that all views are coded at equal spatial resolution. In mixed resolution 3D video coding some of the views are coded at reduced resolution. This paper proposes an improvement for the mode decisions in depth encoding in the mixed resolution scenario. We modify the distortion calculation for rate-distortion optimized depth coding. The proposed solution optimizes the depth data compression with assumption that it will be used not only for view synthesis but also for depth-based super resolution in the post-processing stage. The algorithm is implemented on top of the mixed resolution 3D video encoder based on the 3DV-ATM reference software. Evaluation of the proposed solution, tested under the JCT-3V common test conditions, is done against the mixed resolution MVC+D coding with the view synthesis distortion enabled. The results show 2.64%dBR gain for coded views and 0.64% gain for synthesized views. Michal Joachimiak, Miska M. Hannuksela, Payman Aflaki, Moncef Gabbouj |
ICIP | 4 |
| 2015 | Image interpolation based on non-local geometric similaritiesabstractImage interpolation refers to constructing a high-resolution (HR) image from a low-resolution (LR) image. Traditionally, an HR image can be produced from an observed LR image via the polynomial-based interpolation (bi-linear or bi-cubic interpolations, involving a small number of neighbors around each interpolated position). The advanced interpolation makes use of the so-called “geometric similarity” to design a set of optimal interpolation weighting coefficients. However, better geometric similarities can perhaps be found from a non-local area within the LR source image or even from other but similar images (possibly with higher resolutions). Based on this fact, we propose in this paper a non-local geometric similarity based interpolation scheme to construct HR images. In our proposed method, optimal weighting coefficients are determined by solving a regularized least squares problem which is built upon a number of dual reference patches drawn from the observed LR image and regularized by the variation of directional gradients of the image patch. Experimental results demonstrate that our proposed method offers a remarkable quality improvement, both objectively and subjectively. Shuyuan Zhu, Bing Zeng 0001, Guanghui Liu 0001, Liaoyuan Zeng, Moncef Gabbouj |
ICME | 6 |
| 2015 | On the dynamics of a Recurrent Hopfield NetworkabstractIn this research paper novel real/complex valued recurrent Hopfield Neural Network (RHNN) is proposed. The method of synthesizing the energy landscape of such a network and the experimental investigation of dynamics of Recurrent Hopfield Network is discussed. Parallel modes of operation (other than fully parallel mode) in layered RHNN is proposed. Also, certain potential applications are proposed. Garimella Rama Murthy, Berkay Kicanaoglu, Moncef Gabbouj |
IJCNN | 3 |
| 2015 | On the design of Hopfield Neural Networks: Synthesis of hopfield type associative memoriesabstractIn this research paper, it is proved that it is impossible to design a Hopfield Neural Network with orthogonal stable states ( corners of hypercube ) when the total number of neurons in the network ( dimension of network ) is odd. Also linear algebraic structure of associative memory synthesized by Hopfield is discussed. Using Hadamard matrix of suitable dimension, an algorithm to synthesize real valued Hopfield neural network is discussed. The design of a certain complex Hopfield neural network is addressed and solved. Also, synthesis of real and complex Hopfield type associative memories is discussed. This synthesis enables choice of stable states and the corresponding values of energy function. Garimella Rama Murthy, Moncef Gabbouj |
IJCNN | 2 |
| 2015 | Long-term epileptic EEG classification via 2D mapping and textural features
Kaveh Samiee, Serkan Kiranyaz, Moncef Gabbouj, Tapio Saramäki |
Expert Syst. Appl. | 3 |
| 2015 | Adaptive sampling for compressed sensing based image compression
Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | On the kernel Extreme Learning Machine speedup
Alexandros Iosifidis, Moncef Gabbouj |
Pattern Recognit. Lett. | 2 |
| 2015 | An Efficient Adaptive Binary Range Coder and Its VLSI ArchitectureabstractIn this paper, we propose a new hardware-efficient adaptive binary range coder (ABRC) and its very-large-scale integration (VLSI) architecture. To achieve this, we follow an approach that allows to reduce the bit capacity of the multiplication needed in the interval division part and shows how to avoid the need to use a loop in the renormalization part of ABRC. The probability estimation in the proposed ABRC is based on a lookup table free virtual sliding window. To obtain a higher compression performance, we propose a new adaptive window size selection algorithm. In comparison with an ABRC with a single window, the proposed system provides a faster probability adaptation at the initial encoding/decoding stage, and more accurate probability estimation for very low entropy binary sources. We show that the VLSI architecture of the proposed ABRC attains a throughput of 105.92 MSymbols/s on the FPGA platform, and consumes 18.15 mW for the dynamic part power. In comparison with the state-of-the-art MQ-coder (used in JPEG2000 standard) and the M-coder (used in H.264/Advanced Video Coding and H.265/High Efficiency Video Coding standards), the proposed ABRC architecture provides comparable throughput, reduced memory, and power consumption. Experimental results obtained for a wavelet video codec with JPEG2000-like bit-plane entropy coder show that the proposed ABRC allows to reduce the bit rate by 0.8%-8% in comparison with the MQ-coder and from 1.0%-24.2% in comparison with the M-coder. Eugeniy Belyaev, Kai Liu 0021, Moncef Gabbouj, Yunsong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Constrained Directed Graph Clustering and Segmentation Propagation for Multiple Foregrounds CosegmentationabstractThis paper proposes a new constrained directed graph clustering (DGC) method and segmentation propagation method for the multiple foreground cosegmentation. We solve the multiple object cosegmentation with the perspective of classification and propagation, where the classification is used to obtain the object prior of each class and the propagation is used to propagate the prior to all images. In our method, the DGC method is designed for the classification step, which adds clustering constraints in cosegmentation to prevent the clustering of the noise data. A new clustering criterion such as the strongly connected component search on the graph is introduced. Moreover, a linear time strongly connected component search algorithm is proposed for the fast clustering performance. Then, we extract the object priors from the clusters, and propagate these priors to all the images to obtain the foreground maps, which are used to achieve the final multiple objects extraction. We verify our method on both the cosegmentation and clustering tasks. The experimental results show that the proposed method can achieve larger accuracy compared with both the existing cosegmentation methods and clustering methods. Fanman Meng, Hongliang Li 0001, Shuyuan Zhu, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2015 | Scalable Bit Allocation Between Texture and Depth Views for 3-D Video Streaming Over Heterogeneous NetworksabstractIn the multiview video plus depth (MVD) coding format, both texture and depth views are jointly compressed to represent the 3-D video content. The MVD format enables synthesis of virtual views through depth-image-based rendering; hence, distortion in the texture and depth views affects the quality of the synthesized virtual views. Bit allocation between texture and depth views has been studied with some promising results. However, to the best of our knowledge, most of the existing bit-allocation methods attempt to allocate a fixed amount of total bit rate between texture and depth views; that is, to select appropriate pair of quantization parameters for texture and depth views to maximize the synthesized view quality subject to a fixed total bit rate. In this paper we propose a scalable bit-allocation scheme, where a single ordering of texture and depth packets is derived and used to obtain optimal bit allocation between texture and depth views for any total target rates. In the proposed scheme, both texture and depth views are encoded using the quality scalable coding method; that is, medium grain scalable (MGS) coding of the Scalable Video Coding (SVC) extension of the Advanced Video Coding (H.264/AVC) standard. For varying target total bit rates, optimal bit truncation points for both texture and depth views can be obtained using the proposed scheme. Moreover, we propose to order the enhancement layer packets of the H.264/SVC MGS encoded depth view according to their contribution to the reduction of the synthesized view distortion. On one hand, this improves the depth view packet ordering when considered the rate-distortion performance of synthesized views, which is demonstrated by the experimental results. On the other hand, the information obtained in this step is used to facilitate optimal bit allocation between texture and depth views. Experimental results demonstrate the effectiveness of the proposed scalable bit-allocation scheme for texture and depth views. Jimin Xiao, Miska M. Hannuksela, Tammam Tillo, Moncef Gabbouj, Ce Zhu, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Flexible depth map spatial resolution in depth-enhanced multiview video codingabstractMultiview video plus depth (MVD) has proved to be a promising format enabling various 3D applications. One approach to achieve better MVD compression is to adjust the spatial resolution of depth map based on the content and the application. In this research two schemes are considered: first, multiview video coding accompanied by depth maps to improve the texture coding performance. Second, multiview video plus depth (MVD) coding targeting highest quality for synthesized views. Two algorithms to select the best spatial resolution for each scheme are proposed and the results show 10.8% and 16.5% bitrate reduction compared to the anchor case where depth map resolution is fixed. Payman Aflaki, Miska M. Hannuksela, Moncef Gabbouj |
ICASSP | 3 |
| 2014 | Weighted-prediction-based color gamut scalability extension for the H.265/HEVC video codecabstractColor gamut scalability refers to coding a video in a layered manner where the base and enhancement layers are coded in different color gamut spaces. Color gamut scalability and its relationship with spatial scalability are currently being studied for the scalable extension of HEVC (SHVC) to enable coding of ultra-high definition content having BT.2020 color gamut with 10-bit precision as an enhancement layer and high definition content having BT.709 color gamut with 8-bit precision as the base layer. In this paper, we propose to use the weighted prediction tool of the SHVC standard to map the color gamut of the base layer to the enhancement layer. In addition, we also propose a high-precision bit-depth mapping of the base layer to the enhancement layer that jointly performs upsampling with a bit-depth increase. Simulation results show that these two schemes improve the coding efficiency of the All Intra and Random Access configurations by about 6.8% and 3.6% on average, respectively, compared to a basic scheme where the bit-depth of the base layer is increased by simple bit-shifting. These gains are achieved by imposing no changes to the SHVC standard; hence make the proposed method very useful for practical use-cases as well. Alireza Aminlou, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj |
ICASSP | 4 |
| 2014 | On application of rational Discrete Short Time Fourier Transform in epileptic seizure classificationabstractThis work deals with an adaptive and localized time-frequency representation of time-series signals based on rational functions. The proposed rational Discrete Short Time Fourier Transform (DSTFT) is used for extracting discriminative features in EEG data. We take the advantages of bagging ensemble learning and Alternating Decision Tree (ADTree) classifier to detect the seizure segments in presence of seizure-free segments. The effectiveness of different rational systems is compared with the classical Short Time Fourier Transform (STFT). The comparative study demonstrates that Malmquist-Takenaka rational system outperforms STFT while it can provide a tunable time-frequency representation of the EEG signals and less Mean Square Error (MSE) in the inverse transform. Péter Kovács 0001, Kaveh Samiee, Moncef Gabbouj |
ICASSP | 3 |
| 2014 | Backward compatible enhancement of chroma format in HEVCabstractFirst version of the latest video coding standard, High Efficiency Video Coding (HEVC), only supports coding of video in YUV 4:2:0 chroma format. An extension of the standard that will support other chroma formats is currently under development, however, version 1 decoders will not be able to handle the bitstreams created using this extension. In this paper, we propose a novel method to create scalable bitstreams that involve a backward compatible base layer in 4:2:0 format that can be handled by HEVC version 1 decoders and code additional layers to enhance the chroma resolution. The proposal codes 4:2:0 video in the base layer and the high resolution chroma components as auxiliary pictures as separate enhancement layers. The high resolution chroma components could optionally be predicted from the upsampled 4:2:0 chroma components of base layer. The simulations show that the proposed method achieves scalability with 9.5% coding efficiency penalty on average compared to single layer coding of 4:4:4 video. When compared to simulcast of 4:2:0 and 4:4:4 video, proposed method provides 38% gain on average. Proposed method makes services using high chroma fidelity easier to be deployed, due to the backwards compatibility to existing HEVC implementations with high coding efficiency. Döne Bugdayci Sansli, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj |
ICIP | 4 |
| 2014 | Downward spatially-scalable image reconstruction based on compressed sensingabstractAccording to the compressed sensing (CS) theory, we can sample a sparse signal at a rate that is (much) lower than the required Nyquist rate, while still enabling a nearly exact reconstruction. Image signals are sparse when represented in a certain domain, and because of this, a large number of CS-based image sampling and reconstruction techniques have been developed recently. In this paper, we focus on the design of the downward spatially-scalable image reconstruction from the CS-sampled data. Traditional methods usually reconstruct an image whose size is the same as the original source image and then achieve the downward scalability through sub-sampling. In our proposed method, we unify these two steps into a single one and promise to deliver a much improved quality. Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
ICIP | 4 |
| 2014 | Adaptive sampling for compressed sensing based image compressionabstractThe compressed sensing (CS) theory shows that a sparse signal can be recovered at a sampling rate that is (much) lower than the required Nyquist rate. In practice, many image signals are sparse in a certain domain, and because of this, the CS theory has been successfully applied to the image compression in the past few years. The most popular CS-based image compression scheme is the block-based CS (BCS). In this paper, we focus on the design of an adaptive sampling mechanism for the BCS through a deep analysis of the statistical information of each image block. Specifically, this analysis will be carried out at the encoder side (which needs a few overhead bits) and the decoder side (which requires a feedback to the encoder side), respectively. Two corresponding solutions will be compared carefully in our work. We also present experimental results to show that our proposed adaptive method offers a remarkable quality improvement compared with the traditional BCS schemes. Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
ICME | 3 |
| 2014 | Automatic Object Segmentation by Quantum CutsabstractIn this study, the link between quantum mechanics and graph-cuts is exploited and a novel saliency map generation and salient object segmentation method is proposed based on the ground state solution of a modified Hamiltonian. First, the graph representation of certain quantum mechanical operators is studied. This reveals strong connections with widely used graph-cut algorithms while quantum mechanical constraints exhibit crucial advantages over the existing graph-cut algorithms. Furthermore, concepts such as potential field helps solving a particular singularity problem related to Laplacian matrices. In the proposed approach, the ground state (wave function) corresponding to a sub-atomic particle of a modified Hamiltonian operator corresponds to a particular optimization problem, the solution of which yields the salient object segmentation in a digital image. This approach provides a parameter-free -hence dataset independent-, unsupervised and fully automatic saliency map generation, which outperforms many existing state-of-the-art algorithms. The results of the proposed salient object extraction method exhibit such a promising accuracy that pushes the frontier in this field to the borders of the input-driven processing only - without the use of "object knowledge" aided by long-term human memory and intelligence. Furthermore, with the novel technologies for measuring a quantum wave function, the proposed method has a unique potential: Salient object segmentation in an actual physical setup in nano-scale. Such an unprece-dendent property will not only produce segmentation results instantaneously, but may be a unique opportunity to achieve accurate object segmentation in real-time for the massive visual repositories of today's "Big Data". Çaglar Aytekin, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 3 |
| 2014 | Incremental Learning with Support Vector Data DescriptionabstractDue to the simplicity and firm mathematical foundation, Support Vector Machines (SVMs) have been intensively used to solve classification problems. However, training SVMs on real world large-scale databases is computationally costly and sometimes infeasible when the dataset size is massive and non-stationary. In this paper, we propose an incremental learning approach that greatly reduces the time consumption and memory usage for training SVMs. The proposed method is fully dynamic, which stores only a small fraction of previous training examples whereas the rest can be discarded. It can further handle unseen labels in new training batches. The classification experiments show that the proposed method achieves the same level of classification accuracy as batch learning while the computational cost is significantly reduced, and it can outperform other incremental SVM approaches for the new class problem. Weiyi Xie, Stefan Uhlmann, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 4 |
| 2014 | Hierarchical modeling of F0 contours for voice conversion
Gerard Sanchez, Hanna Silén, Jani Nurminen, Moncef Gabbouj |
INTERSPEECH | 4 |
| 2014 | Content adaptive depth map resampling scheme in multiview video plus depthabstractIn this paper, we propose a content-adaptive pre-processing for downsampling depth map for 3D multiview plus depth video coding. The proposed scheme takes advantage of spatially varying content of depth maps and targets improved encoding performance. The key idea is to divide the video frames into several horizontal and vertical stripes and downsample each of them with an appropriate downsampling ratio. With this approach, the rectangular shape of the frames are preserved, hence the video sequence can be coded with a standard encoder. The downsampling ratios, defined for each Intra period, are determined via a Rate Distortion Optimization process. Simulation results show that the proposed method outperforms the reference in which no downsampling is used by 0.35dB of Bjontegaard delta Peak Signal-to-Noise Ratio (PSNR) and brings up to 31% and average of 18% of Bjontegaard delta bitrate reduction (dBR). Maryam Homayouni, Alireza Aminlou, Payman Aflaki, Moncef Gabbouj |
ISCAS | 4 |
| 2014 | Texture classification using joint statistical representation in space-frequency domain with local quantized patternsabstractDespite its success in texture analysis, Local Binary Pattern (LBP) is operated in the original image space, and it fails to capture deeper pixel interactions to provide a more discriminative description. In this paper, we propose to explore the joint statistical representation in the space-frequency domain with local quantized patterns for texture classification. The proposed method consists of two channels. In each channel, the multi-resolution spatial filters are employed to generate multi-scale spatial maps and the local Fourier transform is subsequently applied to extract local frequency features (spectral maps). The global thresholding is adopted to quantize the spatial and spectral maps into different levels, which are then jointly encoded to built a space-frequency co-occurrence histogram. Finally, the two-channel feature histograms are combined to represent the texture. Experiments on the Outex texture database demonstrate the robustness of our method to image rotation and illumination changes, and our method outperforms the state of the art in terms of the classification accuracy. Tiecheng Song, Hongliang Li 0001, Bing Zeng 0001, Moncef Gabbouj |
ISCAS | 4 |
| 2014 | No reference image quality metric via distortion identification and multi-channel label transferabstractIn this paper, we propose a no reference image quality assessment (NR-IQA) algorithm based on distortion identification (DI) and multi-channel label transfer (LT). First, the distortion type classification is used to obtain the query image's probabilities of belonging to each distortion type. Then, the distortion specific label transfer is implemented in multiple distortion category channels. Based on the hypothesis that the similar images share the similar subjective qualities, the label transfer predicts the subjective quality of the query image by pooling the labels of its k-nearest neighbors (KNN) retrieved from the annotated samples. A weighting average of the multi-channel label transfer's outputs is computed to obtain the final perceptual quality score. The weight is the query image's probability that belongs to the corresponding distortion type. The experimental results show that the proposed method outperforms representative NR-IQA approaches and some full-reference metrics. Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Bing Zeng 0001, Moncef Gabbouj |
ISCAS | 5 |
| 2014 | Adaptive reweighted compressed sensing for image compressionabstractAccording to the compressed sensing (CS) theory, a signal that is sparse in a certain domain can be nearly exactly recovered from a few measurements where the sampling rate is lower than the Nyquist rate. This theory has been successfully applied to the image compression in the past few years as most image signals are highly sparse. In this paper, we apply an adaptive sampling mechanism to the reweighted block-based CS (BCS). The proposed adaptive sampling allocates the measurements to each image block according to the statistical information of the block so as to sample and recover the image more efficiently. Experimental results demonstrate that our adaptive reweighted method offers a very significant quality improvement compared with the traditional BCS schemes, including the non-reweighted and reweighted ones. Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
ISCAS | 3 |
| 2014 | Adaptive Spatial Resolution Selection for Stereoscopic Video Compression with MV-HEVC: A Frequency Based ApproachabstractOne approach for stereoscopic video compression is to down sample the content prior to encoding and up sample it to the original spatial resolution after decoding. In this study it is shown that the ratio by which the content should be rescaled is sequence dependent. Hence, a frequency based method is introduced enabling fast and accurate estimation of the best down sampling ratio for different stereoscopic video clips. It is shown that exploiting this approach can bring 3.38% delta bitrate reduction over five camera-captured sequences. Payman Aflaki, Miska M. Hannuksela, Moncef Gabbouj |
ISM | 3 |
| 2014 | Salient Event Detection in Basketball Mobile VideosabstractModern smartphones have become the most popular means for recording videos. In fact, thanks to their portability, smartphones allow for recording anything and at any moment of our everyday life. One common occasion is represented by sport happenings, where people often record their favourite team or players. Automatic analysis of such videos is important for enabling applications such as automatic organization, browsing and summarization of the content. This paper proposes novel algorithms for the detection of salient events in videos recorded at basketball games. The novel approach consists of jointly analyzing visual data and magnetometer data. The magnetometer data provides information about the horizontal orientation of the camera. The proposed joint analysis allows for a reduced number of false positives and for a reduced computational complexity. The algorithms are tested on data captured during real basketball games. The experimental results clearly show the advantages of the proposed approach. Francesco Cricri, Sujeet Mate, Igor D. D. Curcio, Moncef Gabbouj |
ISM | 4 |
| 2014 | Joint depth and texture filtering targeting MVD compressionabstractTo meet the requirement of current bandwidth and storage facilities, it is required to further decrease the bitrate of 3D video content. This paper presents a scheme to partially filter different regions of the texture views and depth maps taking into account the characteristics of both of them. A combination of the distance of objects from the cameras and the spatial details of texture within the objects defines the location and strength of the smoothing filter that should be applied to that image. A series of subjective tests was conducted to confirm that the perceived quality of the filtered content remains intact and no quality degradations is introduced by the filtering steps. Finally, objective measurements show a Bjontegaard delta bitrate reduction up to 23.77% with an average of 15.2%. Payman Aflaki, Miska M. Hannuksela, Maryam Homayouni, Moncef Gabbouj |
VCIP | 4 |
| 2014 | Improved weighted prediction based color gamut scalability in SHVCabstractOne use case that the scalable extension (SHVC) of the state-of-the-art High Efficiency Video Coding (HEVC) standard aims for is to support Ultra High Definition (UHD) TV broadcast in a backwards compatible way with the existing High Definition (HD) TV broadcast. However, since UHD content typically has higher bit-depth and wider color gamut in addition to increased spatial resolution, the compression efficiency is highly affected by the inter-layer processing applied on the base layer picture. This paper proposes an improvement for the weighted prediction based color gamut scalability to have a better mapping between the color gamuts of the base and enhancement layers. The proposed method aims at capturing the nonlinear characteristics of the color gamut mapping using a piecewise linear model, whose parameters are signaled through weighted prediction mechanism and multiple inter-layer reference pictures. Compared to other existing methods for color gamut mapping in SHVC, such as the 3D Look Up Table (LUT) method, the proposed weighted prediction based approach is less complex, as it does not require any changes to the decoder. The simulation results show up to 3.8% Bjontegaard delta bitrate gain in luma for all intra and 3.0% for random access configurations compared to the existing weighted prediction based scalability method in SHVC. Döne Bugdayci Sansli, Alireza Aminlou, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj |
VCIP | 5 |
| 2014 | Simultaneous 2D and 3D perception for stereoscopic displays based on polarized or active shutter glasses
Payman Aflaki, Miska M. Hannuksela, Hamed Sarbolandi, Moncef Gabbouj |
J. Vis. Commun. Image Represent. | 4 |
| 2014 | Multimodal extraction of events and of information about the recording activity in user generated videosabstractIn this work we propose methods that exploit context sensor data modalities for the task of detecting interesting events and extracting high-level contextual information about the recording activity in user generated videos. Indeed, most camera-enabled electronic devices contain various auxiliary sensors such as accelerometers, compasses, GPS receivers, etc. Data captured by these sensors during the media acquisition have already been used to limit camera degradations such as shake and also to provide some basic tagging information such as the location. However, exploiting the sensor-recordings modality for subsequent higher-level information extraction such as interesting events has been a subject of rather limited research, further constrained to specialized acquisition setups. In this work, we show how these sensor modalities allow inferring information (camera movements, content degradations) about each individual video recording. In addition, we consider a multi-camera scenario, where multiple user generated recordings of a common scene (e.g., music concerts) are available. For this kind of scenarios we jointly analyze these multiple video recordings and their associated sensor modalities in order to extract higher-level semantics of the recorded media: based on the orientation of cameras we identify the region of interest of the recorded scene, by exploiting correlation in the motion of different cameras we detect generic interesting events and estimate their relative position. Furthermore, by analyzing also the audio content captured by multiple users we detect more specific interesting events. We show that the proposed multimodal analysis methods perform well on various recordings obtained in real live music performances. Francesco Cricri, Kostadin Dabov, Igor D. D. Curcio, Sujeet Mate, Moncef Gabbouj |
Multim. Tools Appl. | 5 |
| 2014 | Instance based personalized multi-form image browsing and retrieval
Esin Guldogan, Thomas Olsson 0002, Else Lagerstam, Moncef Gabbouj |
Multim. Tools Appl. | 4 |
| 2014 | Noise-Robust Texture Description Using Local Contrast Patterns via Global MeasuresabstractThis letter presents a noise-robust descriptor by exploring a set of local contrast patterns (LCPs) via global measures for texture classification. To handle image noise, the directed and undirected difference masks are designed to calculate three types of local intensity contrasts: directed, undirected, and maximum difference responses. To describe pixel-wise features, these responses are separately quantized and encoded into specific patterns based on different global measures. These resulting patterns (i.e., LCPs) are jointly encoded to form our final texture representation. Experiments are conducted on the well-known Outex and CUReT databases in the presence of high levels of noise. Compared to many state-of-the-art methods, the proposed descriptor achieves superior texture classification performance while enjoying a compact feature representation. Tiecheng Song, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003, Bing Zeng 0001, Moncef Gabbouj |
IEEE Signal Process. Lett. | 7 |
| 2014 | Differential Coding Using Enhanced Inter-Layer Reference Picture for the Scalable Extension of H.265/HEVC Video CodecabstractDifferential coding methods improve coding efficiency of scalable video codecs by adding the high-frequency component present in the previously coded enhancement layer (EL) pictures to the base layer (BL) picture. This paper proposes a method to enable differential coding in a scalable codec design without affecting the core coding tools, thus allowing a practical implementation to reuse single-layer hardware or software components. This is achieved by creating an additional reference picture called enhanced inter-layer reference (EILR) and inserting it to the EL decoded picture buffer and reference picture lists. An EILR picture is generated by adding differential information to the current inter-layer reference picture. The differential information is calculated using the previously decoded pictures of the BL and EL and the motion information of the BL picture. The proposed method reduces luma total bitrate on average by 2.2% and 2.8% for random access and low-delay test cases, respectively. The improvements are more significant for chroma components with the average bitrate reduction of 6.5%. The measured decoding time increase for a reference software implementation is 16% with negligible overhead on encoding time. Alireza Aminlou, Jani Lainema, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2014 | Perceptual Encryption of H.264 Videos: Embedding Sign-Flips Into the Integer-Based TransformsabstractAn alternative-transforms-based scheme has recently been proposed to achieve perceptual encryption of video signals in which multiple transforms are designed by using different rotation angles at the final stage of the discrete cosine transforms (DCTs) butterfly flow-graph structure. More recently, it is found that a set of more efficient alternative transforms can be derived by introducing sign-flips at the same stage, which is equivalent to an extra rotation angle of π. In this paper, we generalize this sign-flipping technique by randomly embedding sign-flips into all stages of the DCTs butterfly structure so that the encryption space becomes much larger to yield a higher security. We pursue this study for H.264-compatible videos, assuming that the integer DCT of size 4 × 4 is used. First, we follow the separable implementation of the 4 × 4 2-D DCT in which different sign-flipping strategies will be employed along its horizontal and vertical dimensions. Second, we convert the 4 × 4 2-D DCT into a 16-point 1-D butterfly structure so that more sign-flips can be embedded at its various stages. Third, we choose different schemes to pair the node-variables in the 16-point 1-D butterfly structure, thus further enlarging the encryption space. Extensive experiments are conducted to show the performance of these improved encryption schemes and some security analyzes are also presented to confirm their persistence to various attacking strategies. Bing Zeng 0001, Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Moncef Gabbouj |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2014 | Sport Type Classification of Mobile VideosabstractThe recent proliferation of mobile video content has emphasized the need for applications such as automatic organization and automatic editing of videos. These applications could greatly benefit from domain knowledge about the content. However, extracting semantic information from mobile videos is a challenging task, due to their unconstrained nature. We extract domain knowledge about sport events recorded by multiple users, by classifying the sport type into soccer, American football, basketball, tennis, ice-hockey, or volleyball. We adopt a multi-user and multimodal approach, where each user simultaneously captures audio-visual content and auxiliary sensor data (from magnetometers and accelerometers). Firstly, each modality is separately analyzed; then, analysis results are fused for obtaining the sport type. The auxiliary sensor data is used for extracting more discriminative spatio-temporal visual features and efficient camera motion features. The contribution of each modality to the fusion process is adapted according to the quality of the input data. We performed extensive experiments on data collected at public sport events, showing the merits of using different combinations of modalities and fusion methods. The results indicate that analyzing multimodal and multi-user data, coupled with adaptive fusion, improves classification accuracies in most tested cases, up to 95.45%. Francesco Cricri, Mikko Roininen, Jussi Leppänen, Sujeet Mate, Igor D. D. Curcio, Stefan Uhlmann, Moncef Gabbouj |
IEEE Trans. Multim. | 7 |
| 2013 | Intra prediction mode coding for scalable HEVCabstractHigh Efficiency Video Coding (HEVC) standard introduced an increased number of intra prediction directions in order to improve intra prediction performance by efficiently modeling the directional structures found in typical video contents. Efficient coding of intra prediction mode information is realized through a Most Probable Mode (MPM) list approach. In a scalable system, due to high correlation between the layers, utilization of base layer intra prediction mode can improve coding performance. In this paper, we propose a new intra prediction mode coding algorithm for scalable extension of HEVC where only the difference between the intra prediction modes of base and enhancement layers is coded. We provide experimental results and also a comparison of the proposed algorithm with an MPM list based approach where base layer intra prediction mode is added to the list as the most probable mode. Experimental results show BD-rate gains up to 1.1% in 2x spatial scalability and 0.7% in 1.5x scalability for all intra configuration. Döne Bugdayci Sansli, M. Oguz Bici, Kemal Ugur, Moncef Gabbouj |
ICASSP | 4 |
| 2013 | Supervised model training for overlapping sound events based on unsupervised source separationabstractSound event detection is addressed in the presence of overlapping sounds. Unsupervised sound source separation into streams is used as a preprocessing step to minimize the interference of overlapping events. This poses a problem in supervised model training, since there is no knowledge about which separated stream contains the targeted sound source. We propose two iterative approaches based on EM algorithm to select the most likely stream to contain the target sound: one by selecting always the most likely stream and another one by gradually eliminating the most unlikely streams from the training. The approaches were evaluated with a database containing recordings from various contexts, against the baseline system trained without applying stream selection. Both proposed approaches were found to give a reasonable increase of 8 percentage units in the detection accuracy. Toni Heittola, Annamaria Mesaros, Tuomas Virtanen, Moncef Gabbouj |
ICASSP | 4 |
| 2013 | Coding of mixed-resolution multiview video in 3D video applicationabstractThe emerging MVC+D standard specifies the coding of Multiview Video plus Depth (MVD) data for enabling advanced 3D video applications. MVC+D specifications define the coding of all views of MVD at equal spatial resolution and apply a conventional MVC technique for coding the multiview texture and the depth independently. This paper presents a modified MVC+D coding scheme, where only the base view is coded at the original resolution whereas dependent views are coded at reduced resolution. To enable inter-view prediction, the base view is downsampled within the MVC coding loop to provide a relevant reference for dependent views. At the decoder side, the proposed scheme consists of a post-processing scheme which upsamples of the decoded views to their original resolution. The proposed scheme is compared against the original MVC+D scheme and an average of 4% delta bitrate reduction (dBR) in the coded views and 14.5% of dBR in the synthesized views are reported. Payman Aflaki, Wenyi Su, Michal Joachimiak, Dmytro Rusanovskyy, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj |
ICIP | 7 |
| 2013 | Quantum mechanics in computer vision: Automatic object extractionabstractAn automatic object extraction method is proposed exploiting the rich mathematical structure of quantum mechanics. First, a novel segmentation method based on the solutions of Schrödinger's equation is proposed. This powerful segmentation method allows us to model complex objects and inherent structures of edge, shape, and texture information along with the grey-level intensity uniformity, all in a single equation. Due to the large amount of segments extracted with the proposed method, the selection of the object segment is performed by maximizing a regularization energy function based on a recently proposed sub-segment analysis indicating the object boundaries. The results of the proposed automatic object extraction method exhibit such a promising accuracy that pushes the frontier in this field to the borders of the input-driven processing only — without the use of “object knowledge” aided by long-term human memory and intelligence. Çaglar Aytekin, Serkan Kiranyaz, Moncef Gabbouj |
ICIP | 3 |
| 2013 | Efficient video resolution adaptation using scalable H.265/HEVCabstractDynamically changing the spatial resolution in a video conferencing session is useful for seamlessly adapting the bitrate to changing network conditions and for improving the user experience. Similar to earlier standards, the emerging High Efficiency Video Coding (H.265/HEVC) standard does not allow prediction across different resolutions, so an Instantaneous Decoding Refresh (IDR) picture must be sent to reinitialize the stream when a resolution change happens. IDR pictures take significantly more bits compared to predictively coded pictures. Thus, using them for resolution switching significantly reduces coding efficiency and increases the delay. In this paper we propose a method to support efficient adaptive resolution change using the emerging scalable H.265/HEVC standard. The proposed approach utilizes the inter-layer predicted random access pictures at the enhancement layer for resolution switching, instead of IDR pictures. The experimental results show that when the proposed method was used, the bitrate was reduced at the switching point by 34% on average for the tested video sequences. In addition, visual examples are shown demonstrating the improved visual quality with the proposed method. Hoda Roodaki, Kemal Ugur, Miska M. Hannuksela, Moncef Gabbouj |
ICIP | 4 |
| 2013 | Multi-sensor fusion for sport genre classification of user generated mobile videosabstractWe present a robust multimodal approach for classifying the sport genre in videos recorded by mobile phone users at a sport event. In addition to traditional audio-visual content analysis tools, we propose to analyze auxiliary sensor data (electronic compass data and accelerometer data) captured simultaneously with the video recording. By means of machine learning techniques, we build models of visual appearance, camera motion (from auxiliary sensor data) and audio scene, which are used for classifying the data from each modality. The sport genre is obtained by fusing the information provided by the models. We propose to use the quality of each modality as an indication of its reliability. Extensive experiments were performed on real test data collected at public sport events. We provide comparisons on the use of different modality sets and fusion methods. Finally, we show how the proposed methods achieve robust classification even in the considered unconstrained scenarios. Francesco Cricri, Mikko Roininen, Sujeet Mate, Jussi Leppänen, Igor D. D. Curcio, Moncef Gabbouj |
ICME | 6 |
| 2013 | Polarimetric SAR classification using visual color features extracted over pseudo color imagesabstractPolarimetric SAR data have been used extensively for terrain classification applying primitives from various target decompositions as well as texture features. However, there is a source of information that has been neglected so far from polarimetric SAR classification: Color. It is a common practice to visualize polarimetric SAR data by color coding methods and thus it is possible to extract powerful color features from such pseudocolor images. In this paper, we investigate and evaluate discrimination power of color features extracted over various pseudocolor images. Experiments are conducted over the San Francisco Bay region on RADARSAT-2 data by using Support Vector Machines. The classification results show that the additional color features introduce a new level of discrimination and provide noteworthy improvement in classification performance (compared to the traditionally employed polarimetric SAR and texture features) within the application of land use and land cover classification. Stefan Uhlmann, Serkan Kiranyaz, Moncef Gabbouj |
IGARSS | 3 |
| 2013 | Speaker-specific retraining for enhanced compression of unit selection text-to-speech databasesabstractUnit selection based text-to-speech systems can generally obtain high speech quality provided that the database is large enough. In embedded applications, the related memory requirements may be excessive and often the database needs to be both pruned and compressed to fit it into the available memory space. In this paper, we study the topic of database compression. In particular, the focus is on speaker-specific optimization of the quantizers used in the database compression. First, we introduce the simple concept of dynamic quantizer structures, facilitating the use of speaker-specific optimizations by enabling convenient run-time updates. Second, we show that significant memory savings can be obtained through speaker-specific retraining while perfectly maintaining the quantization accuracy, even when the memory required for the additional codebook data is taken into account. Thus, the proposed approach can be considered effective in reducing the conventionally large footprint of unit selection based text-to-speech systems. Index Terms: speech synthesis, unit selection, database compression Jani Nurminen, Hanna Silén, Moncef Gabbouj |
INTERSPEECH | 3 |
| 2013 | Voice conversion for non-parallel datasets using dynamic kernel partial least squares regressionabstractVoice conversion aims at converting speech from one speaker to sound as if it was spoken by another specific speaker. The most popular voice conversion approach based on Gaussian mixture modeling tends to suffer either from model overfitting or oversmoothing. To overcome the shortcomings of the traditional approach, we recently proposed to use dynamic kernel partial least squares (DKPLS) regression in the framework of parallel-data voice conversion. However, the availability of parallel training data from both the source and target speaker is not always guaranteed. In this paper, we extend the DKPLS-based conversion approach for non-parallel data by combining it with a well-known INCA alignment algorithm. The listening test results indicate that high-quality conversion can be achieved with the proposed combination. Furthermore, the performance of two variations of INCA are evaluated with both intra-lingual and cross-lingual data. Index Terms: voice conversion, non-parallel data, kernel partial least squares regression, INCA alignment Hanna Silén, Jani Nurminen, Elina Helander, Moncef Gabbouj |
INTERSPEECH | 4 |
| 2013 | Evaluation of detailed modeling of the LP residual in statistical speech synthesisabstractSpeech parameterization remains an open question in statistical speech synthesis. In our earlier work we have shown that a framework developed originally for highly efficient speech storage can also be successfully applied for voice conversion and concatenative unit selection based speech synthesis. Recently, we have also used the same coding scheme in hybrid-form speech synthesis. In this paper, we further discuss the framework and apply it in statistical speech synthesis, concentrating specifically on the spectral modeling of the linear prediction (LP) residual. Perceptual evaluation demonstrates that the modeling of the spectral details remaining in the residual improves the quality of synthetic speech. Jani Nurminen, Hanna Silén, Elina Helander, Moncef Gabbouj |
ISCAS | 4 |
| 2013 | Flexible Coding Order for 3D video extension of H.265/HEVCabstractThis paper presents a novel Multi-View Video plus Depth (MVD) data coding configuration for the 3D video extension of the High Efficiency Video Coding standard (3D-HEVC). To improve compression efficiency of dependent views, 3D-HEVC utilizes disparity information which is derived from the motion information of the neighboring blocks. However, in the case of MVD data coding, disparity information can be derived directly from the depth data. To utilize this approach, we propose to code the depth maps of the dependent views before their associated texture pictures. Consequently, disparity can be derived from coded depth and can be utilized for coding of texture data. The proposed method improves the compression performance of 3D-HEVC by 0.6% in delta bit rate on average. In addition to this, coded depth maps can be used for Backward View Synthesis Prediction (B-VSP) to further reduce bitrate of dependent texture views. Not only the rate distortion performance is improved, but also the proposed configuration reduces the complexity of decoding by over 8%. Srikanth Gopalakrishna, Miska M. Hannuksela, Moncef Gabbouj |
PCS | 3 |
| 2013 | Nonlinear Depth Map Resampling for Depth-Enhanced 3-D Video CodingabstractDepth-enhanced 3-D video coding includes coding of texture views and associated depth maps. It has been observed that coding of depth map at reduced resolution provides better rate-distortion performance on synthesized views comparing to utilization of full resolution (FR) depth maps in many coding scenarios based on the Advanced Video Coding (H.264/AVC) standard. Conventional techniques for down and upsampling do not take typical characteristics of depth maps, such as distinct edges and smooth regions within depth objects, into account. Hence, more efficient down and upsampling tools, capable of preserving edges better, are needed. In this letter, novel non-linear methods to down and upsample depth maps are presented. Bitrate comparison of synthesized views, including texture and depth map bitstreams, is presented against a conventional linear resampling algorithm. Objective results show an average bitrate reduction of 5.29% and 3.31% for the proposed down and upsampling methods with ratio ½, respectively, comparing to the anchor method. Moreover, a joint utilization of the proposed down and upsampling brings up to 20% and on average 7.35% bitrate reduction. Payman Aflaki, Miska M. Hannuksela, Dmytro Rusanovskyy, Moncef Gabbouj |
IEEE Signal Process. Lett. | 4 |
| 2013 | Multiview-Video-Plus-Depth Coding Based on the Advanced Video Coding StandardabstractThis paper presents a multiview-video-plus-depth coding scheme, which is compatible with the advanced video coding (H.264/AVC) standard and its multiview video coding (MVC) extension. This scheme introduces several encoding and in-loop coding tools for depth and texture video coding, such as depth-based texture motion vector prediction, depth-range-based weighted prediction, joint inter-view depth filtering, and gradual view refresh. The presented coding scheme is submitted to the 3D video coding (3DV) call for proposals (CfP) of the Moving Picture Experts Group standardization committee. When measured with commonly used objective metrics against the MVC anchor, the proposed scheme provides an average bitrate reduction of 26% and 35% for the 3DV CfP test scenarios with two and three views, respectively. The observed bitrate reduction is similar according to an analysis of the results obtained for the subjective tests on the 3DV CfP submissions. Miska M. Hannuksela, Dmytro Rusanovskyy, Wenyi Su, Lulu Chen, Ri Li, Payman Aflaki, Deyan Lan, Michal Joachimiak, Houqiang Li, Moncef Gabbouj |
IEEE Trans. Image Process. | 10 |
| 2013 | Sparse/DCT (S/DCT) Two-Layered Representation of Prediction Residuals for Video CodingabstractIn this paper, we propose a cascaded sparse/DCT (S/DCT) two-layer representation of prediction residuals, and implement this idea on top of the state-of-the-art high efficiency video coding (HEVC) standard. First, a dictionary is adaptively trained to contain featured patterns of residual signals so that a high portion of energy in a structured residual can be efficiently coded via sparse coding. It is observed that the sparse representation alone is less effective in the R-D performance due to the side information overhead at higher bit rates. To overcome this problem, the DCT representation is cascaded at the second stage. It is applied to the remaining signal to improve coding efficiency. The two representations successfully complement each other. It is demonstrated by experimental results that the proposed algorithm outperforms the HEVC reference codec HM5.0 in the Common Test Condition. Je-Won Kang, Moncef Gabbouj, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 2 |
| 2013 | A Low-Complexity Bit-Plane Entropy Coding and Rate Control for 3-D DWT Based Video CodingabstractThis paper is dedicated to fast video coding based on three-dimensional discrete wavelet transform. First, we propose a novel low-complexity bit-plane entropy coding of wavelet subbands based on Levenstein zero-run coder for low entropy contexts and adaptive binary range coder for other contexts. Second, we propose a rate-distortion efficient criterion for skipping 2-D wavelet transforms and entropy encoding based on parent-child subband tree. Finally, we propose one pass rate control which uses virtual buffer concept for adaptive Lagrange multiplier selection. Simulations results show that the proposed video codec has a much lower computational complexity (from 2 to 6 times) for the same quality level compared to the H.264/AVC standard in the low complexity mode. Eugeniy Belyaev, Karen Egiazarian, Moncef Gabbouj |
IEEE Trans. Multim. | 3 |
| 2012 | Local linear transformation for voice conversionabstractMany popular approaches to spectral conversion involve linear transformations determined for particular acoustic classes and compute the converted result as a linear combination between different local transformations in an attempt to ensure a continuous conversion. These methods often produce over-smoothed spectra and parameter tracks. The proposed method computes an individual linear transformation for every feature vector based on a small neighborhood in the acoustic space thus preserving local details. The method effectively reduces the over-smoothing by eliminating undesired contributions from acoustically remote regions. The method is evaluated in listening tests against the well-known Gaussian Mixture Model based conversion, representative of the class of methods involving linear transformations. Perceptual results indicate a clear preference for the proposed scheme. Victor Popa, Hanna Silén, Jani Nurminen, Moncef Gabbouj |
ICASSP | 4 |
| 2012 | An efficient multiplication-free and look-up table-free adaptive binary arithmetic coderabstractIn this paper we propose a novel efficient adaptive binary arithmetic coder which is multiplication-free and requires no look-up tables. To achieve this, we combine the probability estimation based on a virtual sliding window with the approximation of multiplication and the use of simple operations to calculate the next approximation after the encoding of each binary symbol. We show that the proposed algorithm is faster and provides a better compression efficiency compared to the M-coder in the CABAC entropy coding scheme of the H.264/AVC video coding standard. Eugeniy Belyaev, Andrey M. Turlikov, Karen Egiazarian, Moncef Gabbouj |
ICIP | 4 |
| 2012 | A Two-Piece R-D Model for Hybrid Video Coding and Its Application in Fast Mode DecisionabstractThe mode decision process has a significant effect on the quality and complexity of a video encoder. The conventional method that fully codes each macro block for different modes results in the best quality performance, but it suffers from high computational complexity. On the other hand, some other methods ignore the residual part and use the prediction data, or adopt early mode selection approaches in order to reduce the computational cost. These approaches have a negative impact on the coding performance. In this paper, we have used a simple model for the residual coding part and proposed a two-piece R-D model for a macro block. Based on this model, we have introduced a mode decision algorithm that reduces the bit-rate by up to 11.62% at the expense of just 0.5% computational overhead. Alireza Aminlou, Hana Fahim-Hashemi, Mahmoud Reza Hashemi, Moncef Gabbouj, Omid Fatemi |
ICME | 4 |
| 2012 | Ways to Implement Global Variance in Statistical Speech SynthesisabstractHidden Markov model-based speech synthesis is prone to over-smoothing of spectral parameter trajectories. The maximum-likelihood parameter generation favors smooth tracks and the utterance-level variance of each parameter trajectory is significantly reduced compared to the original recordings. This results in muffled speech. To retain the natural variance, statistical global variance modeling has been used in parameter generation. The modeling increases the utterancelevel variance in synthesis, but it is computationally demanding: there is no closed-form solution and an iterative approach is used. In this paper, we analyze the performance of two simple alternative approaches for retaining the natural variance of spectral parameters in synthesis, namely variance scaling and histogram equalization. Both methods apply analytically solvable parameter generation and impose the natural variance afterwards as an efficient post-processing step. Subjective evaluations carried out on English data confirm that the achieved synthesis quality is higher compared to simple post-filtering and similar to the standard global variance modeling. Index Terms: statistical speech synthesis, global variance, variance scaling, histogram equalization Hanna Silén, Elina Helander, Jani Nurminen, Moncef Gabbouj |
INTERSPEECH | 4 |
| 2012 | Complexity analysis of next-generation HEVC decoderabstractThis paper analyzes the complexity of the HEVC video decoder being developed by the JCT-VC community. The HEVC reference decoder HM 3.1 is profiled with Intel VTune on Intel Core 2 Duo processor. The analysis covers both Low Complexity (LC) and High Efficiency (HE) settings for resolutions varying from WQVGA (416 × 240 pixels) up to 1600p (2560 × 1600 pixels). The yielded cycle-accurate results are compared with the respective results of H.264/AVC Baseline Profile (BP) and High Profile (HiP) reference decoders. HEVC offers significant improvement in compression efficiency over H.264/AVC: the average BD-rate saving of LC is around 51% over BP whereas the BD-rate gain of HE is around 45% over HiP. However, the average decoding complexities of LC and HE are increased by 61% and 87% over BP and HiP, respectively. In LC, the most complex functions are motion compensation (MC) and loop filtering (LF) that account on average for 50% and 14% of the decoder complexity. The decoding complexity of HE configuration is on average 42% higher than that of the LC configuration. Majority of the difference is caused by extra LF stages. In HE, the complexities of MC and LF are 37% and 32%, respectively. In practice, a standard 3 GHz dual core processor is expected to be able to decode 1080p HEVC content in real-time. Marko Viitanen, Jarno Vanne, Timo Hämäläinen 0001, Moncef Gabbouj, Jani Lainema |
ISCAS | 4 |
| 2012 | Sensor-Based Analysis of User Generated Video for Multi-camera Video Remixing
Francesco Cricri, Igor D. D. Curcio, Sujeet Mate, Kostadin Dabov, Moncef Gabbouj |
MMM | 5 |
| 2012 | Texture denoising utilized in depth-enhanced multiview video codingabstractIn this study, we applied a locally adaptive filtering in 3D DCT domain to 3 views as a preprocessing stage for encoding. Afterwards, these views were encoded and the decoded views were used by a Depth-Image-Based Rendering algorithm (DIBR) to produce virtual intermediate views. A stereopair at a suitable separation for viewing on a stereoscopic display was selected among the synthesized views. A large-scale subjective assessment of the selected synthesized stereopair was performed. A bitrate reduction of 9% on average and up to 21.8% was achieved with almost no penalties on subjective perceived quality. Payman Aflaki, Dmytro Rusanovskyy, Timo Utriainen, Emilia Pesonen, Miska M. Hannuksela, Satu Jumisko-Pyykkö, Moncef Gabbouj |
PCS | 7 |
| 2012 | Evolutionary RBF classifier for polarimetric SAR images
Turker Ince, Serkan Kiranyaz, Moncef Gabbouj |
Expert Syst. Appl. | 3 |
| 2012 | Dynamic and scalable audio classification by collective network of binary classifiers framework: An evolutionary approach
Serkan Kiranyaz, Toni Mäkinen, Moncef Gabbouj |
Neural Networks | 3 |
| 2012 | Rate adaptation for dynamic adaptive streaming over HTTP in content distribution network
Imed Bouazizi, Miska M. Hannuksela, Moncef Gabbouj |
Signal Process. Image Commun. | 4 |
| 2012 | Voice Conversion Using Dynamic Kernel Partial Least Squares RegressionabstractA drawback of many voice conversion algorithms is that they rely on linear models and/or require a lot of tuning. In addition, many of them ignore the inherent time-dependency between speech features. To address these issues, we propose to use dynamic kernel partial least squares (DKPLS) technique to model nonlinearities as well as to capture the dynamics in the data. The method is based on a kernel transformation of the source features to allow non-linear modeling and concatenation of previous and next frames to model the dynamics. Partial least squares regression is used to find a conversion function that does not overfit to the data. The resulting DKPLS algorithm is a simple and efficient algorithm and does not require massive tuning. Existing statistical methods proposed for voice conversion are able to produce good similarity between the original and the converted target voices but the quality is usually degraded. The experiments conducted on a variety of conversion pairs show that DKPLS, being a statistical method, enables successful identity conversion while achieving a major improvement in the quality scores compared to the state-of-the-art Gaussian mixture-based model. In addition to enabling better spectral feature transformation, quality is further improved when aperiodicity and binary voicing values are converted using DKPLS with auxiliary information from spectral features. Elina Helander, Hanna Silén, Tuomas Virtanen, Moncef Gabbouj |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Zero-Quantized Inter DCT Coefficient Prediction for Real-Time Video CodingabstractSeveral algorithms were proposed to predict the zero-quantized DCT coefficients and reduce the computational complexity of transform and quantization. It is observed that these prediction algorithms achieve good performance for all-zero-quantized DCT blocks. However, the efficiency is much lower for non-all-zero-quantized DCT blocks. This paper proposes an algorithm to improve the prediction efficiency for non-all-zero-quantized DCT blocks. The proposed method extends the prediction to 1-D transforms by developing new Gaussian distribution based thresholds for 1-D transformation. Moreover, the proposed algorithm can perform the prediction on 1-D transforms in both the pixel domain and the transform domain. The prediction for the first stage of 1-D transforms is performed in the pixel domain. However, the second stage of 1-D transforms is performed in the 1-D DCT domain. Because after the first stage of 1-D transforms most energy is concentrated to a few low frequency 1-D DCT coefficients, many transforms in the second stage are skipped. Furthermore, the method fits well the traditional row and column transform structure, and it is more implementation friendly. Simulation results show that the proposed model reduces the complexity of transform and quantization more efficiently than competing techniques. In addition, it is shown that the overall video quality achieved by the proposed algorithm is comparable to the references. Jin Li 0006, Moncef Gabbouj, Jarmo Takala |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Lossless audio hiding method for synchronous audio-video codingabstractThe absence of audio-video synchronization has been one of the most annoying effects in multimedia application. This paper proposes a lossless information hiding method for synchronous audio-video coding. The binary bits of an encoded audio signal are embedded into the DCT domain of the video data to produce synchronous hybrid signal which is further encoded. The decoder can extract the audio bits from the hybrid bit stream, and thus, is able to reconstruct the separate signals. This scheme enables synchronous coding, transmission, storage and playback of a video signal and the associated audio signal. Experimental results show that the proposed method outperforms competing techniques in terms of both the reconstructed video and audio quality. Jin Li 0006, Moncef Gabbouj, Jarmo Takala |
ICASSP | 3 |
| 2011 | Prediction of discrete cosine transformed coefficients in resized pixel blocksabstractA hybrid model was developed to predict the zero quantized discrete cosine transform (ZQDCT) coefficients for intra blocks in our previous work. However, the complicated overhead computations seriously degrade its performance in complexity reduction. This paper proposes a new prediction algorithm with less overhead operations. First, each N × N pixel block at the input of transform is resized to a N/2 × N/2 block. Then, this downsized block is decomposed into a mean value and a residual pixel block. Finally, the N × N 2-D DCT is computed from the mean value and this residual pixel block. Experimental results show that the proposed method reduces more redundant computations than the competing techniques and better real-time performance can be expected. This is particularly suitable for low-power processors with video sequences or great number of images to encode. Jin Li 0006, Moncef Gabbouj, Jarmo Takala, Hexin Chen |
ICASSP | 3 |
| 2011 | Parallel Adaptive HTTP Media StreamingabstractRecent advances in adaptive HTTP streaming of 3GPP have paved the way to the development of media streaming over HTTP. The traditional HTTP streaming relies on a series of request segments and receive segments sequentially. In the current Internet media segments can be delivered through distributed networks. In such a scenario, the traditional series request-receive based adaptive HTTP streaming method is, however, unable to provide optimum streaming since the distributed networks resources are not fully utilized. In this paper, we present a novel parallel adaptive HTTP streaming method. Compared to the traditional series of adaptive HTTP streaming technique, our method a) enables the receiver to request multiple segments in parallel b) provides a solution to maintains a limited number of HTTP sessions for receiving segments in parallel and to determine when to start a new HTTP session to request the next segment c) can adapt media bitrates while receiving previously requested segments. Simulation results show that the proposed parallel adaptive HTTP streaming method outperforms the traditional series of adaptive HTTP streaming with respect to providing a higher playback media quality and decreasing the interruption frequency of media playing. Imed Bouazizi, Moncef Gabbouj |
ICCCN | 3 |
| 2011 | Multi-dimensional evolutionary feature synthesis for content-based image retrievalabstractLow-level features (also called descriptors) play a central role in content-based image retrieval (CBIR) systems. Features are various types of information extracted from the content and represent some of its characteristics or signatures. However, especially the (low-level) features, which can be extracted automatically usually lack the discrimination power needed for accurate description of the image content and may lead to a poor retrieval performance. In order to efficiently address this problem, in this paper we propose a multi- dimensional evolutionary feature synthesis technique, which seeks for the optimal linear and non-linear operators so as to synthesize highly discriminative set of features in an optimal dimension. The optimality therein is sought by the multi-dimensional particle swarm optimization method along with the fractional global-best formation technique. Clustering and CBIR experiments where the proposed feature synthesizer is evolved using only the minority of the image database, demonstrate a significant performance improvement and exhibit a major discrimination between the features of different classes. Serkan Kiranyaz, Jenni Raitoharju, Turker Ince, Moncef Gabbouj |
ICIP | 4 |
| 2011 | Incremental evolution of collective network of binary classifier for polarimetric SAR image classificationabstractIn this paper, we propose a dedicated application of collective network of binary classifiers (CNBC) to address the problem of incremental learning, which occurs by introducing new SAR terrain classes. Furthermore, another major goal is to achieve a high classification performance over multiple SAR images even though the training data may not be entirely accurate. The CNBC in principle adopts a “Divide and Conquer” type approach by allocating an individual network of binary classifiers (NBCs) to discriminate each SAR terrain class among others and performing evolutionary search to find the optimal binary classifier (BC) in each NBC. Such design further allows dynamic SAR class and feature scalability in such a way that the CNBC can gradually adapt its internal topology to new features and classes with minimal effort. Experiments visually demonstrate the classification accuracy and efficiency of the proposed system over eight fully polarimetric NASA/JPL AIRSAR data sets. Stefan Uhlmann, Serkan Kiranyaz, Moncef Gabbouj, Turker Ince |
ICIP | 3 |
| 2011 | Fast motion estimation with dual search window for stereo 3d video encodingabstractStereoscopic 3D video is becoming a reality in many application areas, ranging from high quality entertainment to mobile video services. Due to the need to process two views, the complexity of 3D video applications is significantly higher than traditional 2D counterparts. In order to enable real-time 3D video services in mobile devices, this paper proposes a novel algorithm which reduces the complexity of stereo video encoding with improvement of coding efficiency. A novel search window center prediction method is proposed that exploits the correlation between two views. Experimental results show that the average encoding time of the second view can be decreased by 80% with an increase in coding efficiency of up to 2%. The state-of-art fast motion estimation methods for stereoscopic 3D video encoding show coding efficiency decrease, whereas proposed method achieves the speed-up with increase in coding efficiency, making it suitable for high quality 3D video applications. Michal Joachimiak, Kemal Ugur, Jani Lainema, Moncef Gabbouj |
ICME | 4 |
| 2011 | Segment duration for rate adaptation of adaptive HTTP streamingabstractRecently, 3GPP packet-switched streaming (PSS) specified adaptive HTTP streaming (AHS). The client requests series of media segments and adapt the bitrates of segment to varying network resources. The current state-of-the-art rate adaptation of AHS estimates the end-to-end network capacity using the previous segment reception such as segment fetching time. However, the accuracy and the speed of such rate adaptation highly depend on the segments duration. This paper presents a novel segment duration determination method to provide accurate and fast rate adaptation of AHS. The segment duration is estimated as the minimum duration to produce the smoothed HTTP/TCP rate which represents the current network capacity. Rate adaptation method using the determined segment duration is presented. To prevent buffer draining-up, the segment duration is further restrained as fine-grained duration. Simulation results show that the proposed segment duration enables the rate adaptation algorithm to increase achievable media bitrates and reduce the play-back interruption compared with the state-of-the-art rate adaptation method of AHS. Imed Bouazizi, Moncef Gabbouj |
ICME | 3 |
| 2011 | Prediction Signal Aided Spatially Varying TransformabstractSpatially Varying Transform (SVT) is a technique introduced earlier to improve the coding efficiency of video coders [1][2]. SVT allows the position of the transform block within the macroblock to vary in order to better localize the underlying residual signal. The coding gains of SVT come with increased encoding complexity due to the additional need in the encoder to search for the best Location Parameter (LP) which indicates the position of the transform. In this paper, a new technique called Prediction Signal Aided Spatially Varying Transform (PSASVT) is proposed that utilizes the gradient of prediction signal to eliminate the unlikely LPs. As the number of candidate LPs is reduced, a smaller number of LPs are searched by encoder, which reduces the encoding complexity. In addition, less overhead bits are needed to code the selected LP and thus the coding efficiency can be improved. Experimental results show that the number of LPs to be tested in RDO is reduced on average by more than 20%. This reduction in encoding complexity is achieved with a slight increase in coding efficiency, as the number of candidate LPs is reduced. The decoding complexity increase is only a little. Cixun Zhang, Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
ICME | 5 |
| 2011 | Prediction of Voice Aperiodicity Based on Spectral Representations in HMM Speech SynthesisabstractIn hidden Markov model-based speech synthesis, speech is typically parameterized using source-filter decomposition. A widely used analysis/synthesis framework, STRAIGHT, decomposes the speech waveform into a framewise spectral envelope and a mixed mode excitation signal. Inclusion of an aperiodicity measure in the model enables synthesis also for signals that are not purely voiced or unvoiced. In the traditional approach employing hidden Markov modeling and decision tree-based clustering, the connection between speech spectrum and aperiodicities is not taken into account. In this paper, we take advantage of this dependency and predict voice aperiodicities afterwards based on synthetic spectral representations. The evaluations carried out for English data confirm that the proposed approach is able to provide prediction accuracy that is comparable to the traditional approach. Index Terms: aperiodicity prediction, hidden Markov model, speech synthesis Hanna Silén, Elina Helander, Moncef Gabbouj |
INTERSPEECH | 3 |
| 2011 | Multimodal Event Detection in User Generated VideosabstractNowadays most camera-enabled electronic devices contain various auxiliary sensors such as accelerometers, gyroscopes, compasses, GPS receivers, etc. These sensors are often used during the media acquisition to limit camera degradations such as shake and also to provide some basic tagging information such as the location used in geo-tagging. Surprisingly, exploiting the sensor-recordings modality for high-level event detection has been a subject of rather limited research, further constrained to highly specialized acquisition setups. In this work, we show how these sensor modalities, alone or in combination with content-based analysis, allow inferring information about the video content. In addition, we consider a multi-camera scenario, where multiple user generated recordings of a common scene (e.g., music concerts, public events) are available. In order to understand some higher-level semantics of the recorded media, we jointly analyze the individual video recordings and sensor measurements of the multiple users. The detected semantics include generic interesting events and some more specific events. The detection exploits correlations in the camera motion and in the audio content of multiple users. We show that the proposed multimodal analysis methods perform well on various recordings obtained in real live music performances. Francesco Cricri, Kostadin Dabov, Igor D. D. Curcio, Sujeet Mate, Moncef Gabbouj |
ISM | 5 |
| 2011 | Rate adaptation for adaptive HTTP streamingabstractRecently, HTTP has been widely used for the delivery of real-time multimedia content over the Internet, such as in video streaming applications. To combat the varying network resources of the Internet, rate adaptation is used to adapt the transmission rate to the varying network capacity. A key research problem of rate adaptation is to identify network congestion early enough and to probe the spare network capacity. In adaptive HTTP streaming, this problem becomes challenging because of the difficulties in differentiating between the short-term throughput variations, incurred by the TCP congestion control, and the throughput changes due to more persistent bandwidth changes. Imed Bouazizi, Moncef Gabbouj |
MMSys | 3 |
| 2011 | Multi-dimensional particle swarm optimization in dynamic environments
Serkan Kiranyaz, Jenni Raitoharju, Moncef Gabbouj |
Expert Syst. Appl. | 3 |
| 2011 | Personalized long-term ECG classification: A systematic approach
Serkan Kiranyaz, Turker Ince, Jenni Raitoharju, Moncef Gabbouj |
Expert Syst. Appl. | 4 |
| 2011 | A generic content-based image retrieval framework for mobile devices
Iftikhar Ahmad 0001, Moncef Gabbouj |
Multim. Tools Appl. | 2 |
| 2011 | Video Coding Using Spatially Varying TransformabstractIn this paper, a novel algorithm called spatially varying transform (SVT) is proposed to improve the coding efficiency of video coders. SVT enables video coders to vary the position of the transform block, unlike state-of-art video codecs where the position of the transform block is fixed. In addition to changing the position of the transform block, the size of the transform can also be varied within the SVT framework, to better localize the prediction error so that the underlying correlations are better exploited. It is shown in this paper that by varying the position of the transform block and its size, characteristics of prediction error are better localized, and the coding efficiency is thus improved. The proposed algorithm is implemented and studied in the H.264/AVC framework. We show that the proposed algorithm achieves 5.85% bitrate reduction compared to H.264/AVC on average over a wide range of test set. Gains become more significant at medium to high bitrates for most tested sequences and the bitrate reduction may reach 13.50%, which makes the proposed algorithm very suitable for future video coding solutions focusing on high fidelity video applications. The gain in coding efficiency is achieved with a similar decoding complexity which makes the proposed algorithm easy to be incorporated in video codecs. However, the encoding complexity of SVT can be relatively high because of the need to perform a number of rate distortion optimization (RDO) steps to select the best location parameter (LP), which indicates the position of the transform. In this paper, a novel low complexity algorithm is also proposed, operating on a macroblock and a block level, to reduce the encoding complexity of SVT. Experimental results show that the proposed low complexity algorithm can reduce the number of LPs to be tested in RDO by about 80% with only a marginal penalty in the coding efficiency. Cixun Zhang, Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2010 | Dynamic Data Clustering Using Stochastic Approximation Driven Multi-Dimensional Particle Swarm Optimization
Serkan Kiranyaz, Turker Ince, Moncef Gabbouj |
EvoApplications (1) | 3 |
| 2010 | Advanced Rate Adaption for Unicast Streaming of Scalable VideoabstractRecent advances in scalable video coding have paved the way for the development of flexible and adaptive media streaming applications. In order to support rate adaptation, 3GPP Packet-Switched Streaming Service (PSS) has specified a feedback mechanism that includes transmitting information about the status of the client buffer. This paper presents a novel Multiple Virtual Client Buffer Feedback (MVCBF) mechanism, which includes information about multiple media times each of which corresponding to a different set of sub-streams in scalable media streaming. Moreover, a new rate adaptation method is presented that a) assigns different requirements for the buffering time to the different sets of sub-streams, b) observes the media time variation at sub-stream level, and c) efficiently changes the operation point to maintain the required buffering time in priority order. Simulation results show that the proposed MVCBF-based rate adaptation algorithm outperforms the traditional PSS compliant rate adaptation method in the overall throughput by quickly adapting the operation point to the varying network resources and avoiding unnecessary bouncing between operation points. Imed Bouazizi, Moncef Gabbouj |
ICC | 3 |
| 2010 | Subjective study on compressed asymmetric stereoscopic videoabstractAsymmetric stereoscopic video coding takes advantage of the binocular suppression of the human vision by representing one of the views with a lower quality. This paper describes a subjective quality test with asymmetric stereoscopic video. Different options for achieving compressed mixed-quality and mixed-resolution asymmetric stereo video were studied and compared to symmetric stereo video. The bitstreams for different coding arrangements were simulcast-coded according to the Advanced Video Coding (H.264/AVC) standard. The results showed that in most cases, resolution-asymmetric stereo video with the downsampling ratio of 1/2 along both coordinate axes provided similar quality as symmetric and quality-asymmetric full-resolution stereo video. These results were achieved under same bitrate constrain while the processing complexity decreased considerably. Moreover, in all test cases, the symmetric and mixed-quality full-resolution stereoscopic video bitstreams resulted in a similar quality at the same bitrates. Payman Aflaki, Miska M. Hannuksela, Jukka Häkkinen, Paul Lindroos, Moncef Gabbouj |
ICIP | 5 |
| 2010 | Congestion-aware transmission rate control using Medium Grain Scalability of Scalable Video CodingabstractIn packet-oriented networks, packet losses occur mainly due to queue overflows in congested network elements. An increase of the packet rate therefore raises the likelihood of congestion and expected number of lost packets. However, an increase of the packet size has typically a negligible impact on the packet loss rate or congestion in wired packet-switched networks. Scalable Video Coding (SVC), an extension of the Advanced Video Coding (H.264/AVC), provides different types of scalability, one of which is Medium Grain Quality Scalability (MGS). When MGS is in use, layer representations can be pruned unevenly without affecting the decoding of the remaining bitstream. The article proposes a packetization algorithm to improve the quality of the reconstructed video stream while the expected packet loss rate remains unchanged compared to conventional packetization. The algorithm is based on appending MGS enhancement layer data into conventionally generated packet payloads until the Maximum Transmission Unit (MTU) is reached. The simulation results show that the proposed algorithm provides 0.3 to 0.5 dB gain in average luma Peak Signal-to-Noise Ratio when compared with a conventional packetization method while the packet rate remains unchanged. Miska M. Hannuksela, Haibo Zhu, Houqiang Li, Moncef Gabbouj |
ICIP | 4 |
| 2010 | Network of evolutionary binary classifiers for classification and retrieval in macroinvertebrate databasesabstractIn this paper, we focus on advanced classification and data retrieval schemes that are instrumental when processing large taxonomical image datasets. With large number of classes, classification and an efficient retrieval of a particular benthic macroinvertebrate image within a dataset will surely pose a severe problem. To address this, we propose a novel network of evolutionary binary classifiers, which is scalable, dynamically adaptable and highly accurate for the classification and retrieval of large biological species-image datasets. The classification and retrieval results for the macroinvertebrate test data attain taxonomic accuracy that equals and even surpasses that of an average expert. Our findings are encouraging for aquatic biomonitoring where cost intensity of sample analysis currently poses a bottleneck for routine biomonitoring. Serkan Kiranyaz, Moncef Gabbouj, Jenni Raitoharju, Turker Ince, Kristian Meissner |
ICIP | 2 |
| 2010 | Depth-level-adaptive view synthesis for 3D videoabstractIn the multiview video plus depth (MVD) representation for 3D video, a depth map sequence is coded for each view. In the decoding end, a view synthesis algorithm is used to generate virtual views from depth map sequences. Many of the known view synthesis algorithms introduce rendering artifacts especially at object boundaries. In this paper, a depth-level-adaptive view synthesis algorithm is presented to reduce the amount of artifacts and to improve the quality of the synthesized images. The proposed algorithm introduces awareness of the depth level so that no pixel value in the synthesized image is derived from pixels of more than one depth level. Improvements on objective quality of the synthesized views were achieved in five out of eight test cases, while the subjective quality of the proposed method was similar to or better than that of the view synthesis method used by Moving Picture Experts Group (MPEG). Ying Chen 0011, Weixing Wan, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj |
ICME | 6 |
| 2010 | LDPC FEC code extension for unequal error protection in 2nd generation DVB systemsabstractOne of the envisioned advantages of scalable video coding is its inherent suitability for achieving unequal error protection (UEP). UEP can be effectively used for graceful quality degradation under harsh network conditions. The Digital Video Broadcasting (DVB) organization recently introduced second generations of broadcast transmission technologies like DVB-T2 and DVB-S2. These technologies, due to the newly introduced and advanced tools, ensure better quality of service compared to their respective first generation counterparts. Among the important tools that directly affect the quality of service positively in these emerging technologies is the new physical layer-chained forward error correction coding. In both cases, the chained codes comprise of a Bose-Chaudhuri-Hocquenghem (BCH) code, which functions as the outer code, followed by a Low-Density Parity Check (LDPC) code as the inner code. To allow for improved graceful quality degradation in the second generation DVB broadcast technologies, this article proposes a novel method to extend deployed LDPC codes. This code extension enables the implementation of UEP schemes in an effective manner in these DVB broadcast systems. Furthermore, the proposed solution is backward compatible to legacy receivers. The performance evaluation of the proposed method is supported by results, obtained through simulations. Lukasz Kondrad, Imed Bouazizi, Moncef Gabbouj |
ICME | 3 |
| 2010 | Multi-buffer based congestion control for multicast streaming of scalable videoabstractReceiver driven layered multicast streaming provides an attractive solution for transmitting the same video data to multiple receivers while accounting for the heterogeneity in the network resources and device capabilities. Traditionally, congestion control methods in receiver-driven layered multicast have used packet loss as a measure to detect congestion. However, given that wireless networks are characterized by higher packet loss and varying throughputs, packet-loss based congestion control results in inappropriate behavior of those algorithms, and ultimately in sub-optimal usage of the available network resources. This paper presents a novel multicast congestion control algorithm named Layered Virtual Client Buffer (LVCB)-based receiver-driven multicast for multicasting of scalable video. The proposed LVCB technique tracks the media time for each layer currently present in the receiver buffer. The proposed multicast congestion control method reacts to variations in the media time for each LVCB by dynamically joining/leaving multicast groups, in order to adapt the subscription level to the varying network resources. Furthermore, the presented algorithm solves the problem of mutual affection between receivers without exchange of information about subscription levels. The simulation results show the suitability of the proposed method in wireless as well as wired scenarios. Imed Bouazizi, Moncef Gabbouj |
ICME | 3 |
| 2010 | Classification of Polarimetric SAR Images Using Evolutionary RBF NetworksabstractThis paper proposes an evolutionary RBF network classifier for polar metric synthetic aperture radar ( SAR) images. The proposed feature extraction process utilizes the full covariance matrix, the gray level co-occurrence matrix (GLCM) based texture features, and the backscattering power (Span) combined with the H/α/A decomposition, which are projected onto a lower dimensional feature space using principal component analysis. An experimental study is performed using the fully polar metric San Francisco Bay data set acquired by the NASA/Jet Propulsion Laboratory Airborne SAR (AIRSAR) at L-band to evaluate the performance of the proposed classifier. Classification results (in terms of confusion matrix, overall accuracy and classification map) compared to the Wish art and a recent NN-based classifiers demonstrate the effectiveness of the proposed algorithm. Turker Ince, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 3 |
| 2010 | Maximum a posteriori voice conversion using sequential monte carlo methodsabstractMany voice conversion algorithms are based on frame-wise mapping from source features into target features. This ignores the inherent temporal continuity that is present in speech and can degrade the subjective quality. In this paper, we propose to optimize the speech feature sequence after a frame-based conversion algorithm has been applied. In particular, we select the sequence of speech features through the minimization of a cost functionthatinvolvesboththeconversionerrorandthesmoothness of the sequence. The estimation problem is solved using sequential MonteCarlo methods. Both subjectiveand objective results show the effectiveness of the method. Index Terms: voice conversion, maximum a posteriori,Viterbi algorithm, smoothing, particlefilter Elina Helander, Hanna Silén, Joaquín Míguez, Moncef Gabbouj |
INTERSPEECH | 4 |
| 2010 | Using robust viterbi algorithm and HMM-modeling in unit selection TTS to replace units of poor qualityabstractIn hidden Markov model-based unit selection synthesis, the benefits of both unit selection and statistical parametric speech synthesis are combined. However, conventional Viterbi algorithm is forced to do a selection also when no suitable units are available. This can drift the search and decrease the overall quality. Consequently, we propose to use robust Viterbi algorithm that can simultaneously detect bad units and select the best sequence. The unsuitable units are replaced using hidden Markov model-based synthesis. Evaluations indicate that the use of robust Viterbi algorithm combined with unit replacement increases the quality compared to the traditional algorithm. Hanna Silén, Elina Helander, Jani Nurminen, Konsta Koppinen, Moncef Gabbouj |
INTERSPEECH | 5 |
| 2010 | Efficient SIMD-based implementation of adaptive filterabstractDirectional Adaptive Interpolation Filtering (DAIF) is a novel interpolation technique that was proposed recently for hybrid video coding. It was reported, that this technique outperforms the standard H.264/AVC interpolation in terms of coding gain whereas requiring smaller number of arithmetic operations. In this publication we present an optimized implementation of DAIF on a modern computing platform exploiting the Single Instruction Multiple Data (SIMD) parallelism. In addition, we provide a complexity analysis in which the computational complexity is estimated as number of clock cycles per output sample. Proposed SIMD-based implementation of DAIF has lower or comparable interpolation complexity, compared to the highly optimized SIMD-based implementation of the H.264/AVC interpolations. Considering significantly better coding gain provided by DAIF, we believe this approach will play a significant role in future video coding standards. Antti Hallapuro, Dmytro Rusanovskyy, Kemal Ugur, Jani Lainema, Moncef Gabbouj |
ISCAS | 5 |
| 2010 | Time-variable camera separation for compression of stereoscopic videoabstractThis paper presents a hypothesis that stereoscopic perception requires a short adjustment period after a scene change before it is fully effective. A compression method based on this hypothesis is proposed - instead of coding pictures from the left and right views conventionally, a view in the middle of the left and right view is coded for a limited period after a scene change. The coded middle view can be utilized in two alternative ways in rendering. First, it can be rendered as such, which causes an abrupt change from conventional monoscopic video to stereoscopic video. Second, the layered depth video (LDV) coding scheme can be used to associate depth, background texture, and background depth to the middle view, enabling view synthesis and gradual view disparity increase in rendering. Subjective experiments were conducted to evaluate and validate the presented hypothesis and compare the two rendering methods. The results indicate that when the maximum disparity between the left and right views was relatively small, the presented time-variable camera separation method was imperceptible. A compression gain, the magnitude of which depended on the scene duration, was achieved with half of the sequences having a suitable disparity for the presented coding method. Maosheng Ji, Miska M. Hannuksela, Moncef Gabbouj, Houqiang Li |
VCIP | 3 |
| 2010 | Evaluation of global and local training techniques over feed-forward neural network architecture spaces for computer-aided medical diagnosis
Turker Ince, Serkan Kiranyaz, Jenni Raitoharju, Moncef Gabbouj |
Expert Syst. Appl. | 4 |
| 2010 | Perceptual color descriptor based on spatial distribution: A top-down approach
Serkan Kiranyaz, Murat Birinci, Moncef Gabbouj |
Image Vis. Comput. | 3 |
| 2010 | Perceptual-based quality assessment for audio-visual services: A survey
Junyong You, Ulrich Reiter, Miska M. Hannuksela, Moncef Gabbouj, Andrew Perkis |
Signal Process. Image Commun. | 4 |
| 2010 | Voice Conversion Using Partial Least Squares RegressionabstractVoice conversion can be formulated as finding a mapping function which transforms the features of the source speaker to those of the target speaker. Gaussian mixture model (GMM)-based conversion is commonly used, but it is subject to overfitting. In this paper, we propose to use partial least squares (PLS)-based transforms in voice conversion. To prevent overfitting, the degrees of freedom in the mapping can be controlled by choosing a suitable number of components. We propose a technique to combine PLS with GMMs, enabling the use of multiple local linear mappings. To further improve the perceptual quality of the mapping where rapid transitions between GMM components produce audible artefacts, we propose to low-pass filter the component posterior probabilities. The conducted experiments show that the proposed technique results in better subjective and objective quality than the baseline joint density GMM approach. In speech quality conversion preference tests, the proposed method achieved 67% preference score against the smoothed joint density GMM method and 84% preference score against the unsmoothed joint density GMM method. In objective tests the proposed method produced a lower Mel-cepstral distortion than the reference methods. Elina Helander, Tuomas Virtanen, Jani Nurminen, Moncef Gabbouj |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | Fractional Particle Swarm Optimization in Multidimensional Search SpaceabstractIn this paper, we propose two novel techniques, which successfully address several major problems in the field of particle swarm optimization (PSO) and promise a significant breakthrough over complex multimodal optimization problems at high dimensions. The first one, which is the so-called multidimensional (MD) PSO, re-forms the native structure of swarm particles in such a way that they can make interdimensional passes with a dedicated dimensional PSO process. Therefore, in an MD search space, where the optimum dimension is unknown, swarm particles can seek both positional and dimensional optima. This eventually removes the necessity of setting a fixed dimension a priori, which is a common drawback for the family of swarm optimizers. Nevertheless, MD PSO is still susceptible to premature convergences due to lack of divergence. Among many PSO variants in the literature, none yields a robust solution, particularly over multimodal complex problems at high dimensions. To address this problem, we propose the fractional global best formation (FGBF) technique, which basically collects all the best dimensional components and fractionally creates an artificial global best (aGB) particle that has the potential to be a better "guide" than the PSO's native gbest particle. This way, the potential diversity that is present among the dimensions of swarm particles can be efficiently used within the aGB particle. We investigated both individual and mutual applications of the proposed techniques over the following two well-known domains: 1) nonlinear function minimization and 2) data clustering. An extensive set of experiments shows that in both application domains, MD PSO with FGBF exhibits an impressive speed gain and converges to the global optima at the true dimension regardless of the search space dimension, swarm size, and the complexity of the problem. Serkan Kiranyaz, Turker Ince, Alper Yildirim, Moncef Gabbouj |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2009 | Video coding using Variable Block-Size Spatially Varying TransformsabstractIn our previous work, we introduced Spatially Varying Transforms (SVT) for video coding, where the location of the transform block within the macroblock is not fixed but varying. In this paper, we extend this concept and present a novel method, called Variable Block-size Spatially Varying Transforms (VBSVT). VBSVT utilizes Variable Block-size Transforms (VBT) in the SVT framework, and is shown to be more preferable for coding prediction error with different characteristics than fixed block-size SVT and also the standard methods that use fixed or adaptive block sizes at fixed spatial locations. In addition, VBSVT has similar decoding complexity with fixed block-size SVT and lower decoding complexity compared to standard methods as only a portion of the prediction error needs to be decoded. Experimental results show that, VBSVT achieves 4.1% gain over H.264/AVC on average over a wide range of test set. Gains become more significant at high quality levels and go up to 13.5%, which makes the proposed algorithm very suitable for future video coding solutions focusing on high fidelity applications. Cixun Zhang, Kemal Ugur, Jani Lainema, Moncef Gabbouj |
ICASSP | 4 |
| 2009 | Adaptive interpolation with flexible filter structures for video codingabstractTwo novel algorithms are proposed for improving the coding efficiency of adaptive interpolation schemes for video codecs, without increasing implementation complexity. Proposed algorithms utilize two different filter structures with equal tap length, but with complementary frequency responses. Depending on the content being coded, encoder selects which one of the two filter structure is optimal and signals this information to the decoder. In addition, the symmetry of filters is not pre-defined but is flexible. Encoder selects the optimal filter symmetry depending on the coding rate and the content and signals this information to the decoder. Experimental results show, that proposed improvements bring up to 7% of bit-rate reduction at high bit-rate over conventional adaptive interpolation. When compared to H.264/AVC, average gain over the test set is 11%. Coding efficiency is improved without increasing the complexity, thus proposed algorithms are suitable for mobile multimedia use-cases, where the computational resources are very limited. Dmytro Rusanovskyy, Kemal Ugur, Moncef Gabbouj |
ICIP | 3 |
| 2009 | An objective video quality metric based on spatiotemporal distortionabstractThis paper proposes an objective video quality metric based on an analysis of spatial and temporal distortions. Spatial quality features extracted from the spatiotemporal region of reference and distorted videos are used to express the spatial distortion. Temporal distortion, caused by frame freezing resulting from a packet loss, is derived from the spatial distortion before and after the frozen frames. The overall quality is predicted according to the weighted combination of qualities over all the temporal regions. The experimental results with respect to the subjective measurements demonstrate the fast computation and promising performance of the proposed model compared with existing methods. Junyong You, Miska M. Hannuksela, Moncef Gabbouj |
ICIP | 3 |
| 2009 | A novel technique for voice conversion based on style and content decomposition with bilinear modelsabstractThis paper presents a novel technique for voice conversion by solving a two-factor task using bilinear models. The spectral content of the speech represented as line spectral frequencies is separated into so-called style and content parameterizations using a framework proposed in [1]. This formulation of the voice conversion problem in terms of style and content offers a flexible representation of factor interactions and facilitates the use of efficient training algorithms based on singular value decomposition and expectation maximization. Promising results in a comparison with the traditional Gaussian mixture model based method indicate increased robustness with small training sets. 1. Victor Popa, Jani Nurminen, Moncef Gabbouj |
INTERSPEECH | 3 |
| 2009 | Parameterization of vocal fry in HMM-based speech synthesisabstractHMM-based speech synthesis offers a way to generate speech with different voice qualities. However, sometimes databases contain certain inherent voice qualities that need to be parametrized properly. One example of this is vocal fry typically occurring at the end of utterances. A popular mixed excitation vocoder for HMM-based speech synthesis is STRAIGHT. The standard STRAIGHT is optimized for modal voices and may not produce high quality with other voice types. Fortunately, due to the flexibility of STRAIGHT, different F0 and aperiodicity measures can be used in the synthesis without any inherent degradations in speech quality. We have replaced the STRAIGHT excitation with a representation based on a robust F0 measure and a carefully determined two-band voicing. According to our analysis-synthesis experiments, the new parameterization can improve the speech quality. In HMM-based speech synthesis, the quality is significantly improved especially due to the better modeling of vocal fry. Index Terms: speech synthesis, hidden Markov models, vocal fry, mixed excitation, STRAIGHT Hanna Silén, Elina Helander, Jani Nurminen, Moncef Gabbouj |
INTERSPEECH | 4 |
| 2009 | Regionally Adaptive Filtering for Asymmetric Stereoscopic Video CodingabstractIn asymmetric stereoscopic video coding, one view can be coded in a lower resolution of the other. In this scenario, stereoscopic video can be compressed with only moderately increased bandwidth and complexity compared to 2D monoview video coding. The subjective quality degradation of this scenario can be negligible compared to coding two views with original resolution. The low-resolution view can be predicted from the high-resolution view to achieve higher coding efficiency. In this paper, a regionally adaptive filtering algorithm is proposed to generate a predictor of a macroblock (MB) or MB partition of the low-resolution view from the high-resolution view. Different filters are applied for different picture regions. Disparity motion matching and clustering are applied in the encoder for generation of regionally adaptive filters. Simulation results show that the proposed algorithm results in up to 27% bit-rate saving compared with methods without adaptive filtering. Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela |
ISCAS | 3 |
| 2009 | Joint Texture and Depth Map Video Coding based on the Scalable Extension of H.264/AVCabstractDepth-Image-Based Rendering (DIBR) is widely used for view synthesis in 3D video applications. Compared with traditional 2D video applications, both the texture video and its associated depth map are required for transmission in a communication system that supports DIBR. To efficiently utilize limited bandwidth, coding algorithms, e.g. the Advanced Video Coding (H.264/AVC) standard, can be adopted to compress the depth map using the 4:0:0 chroma sampling format. However, when the correlation between texture video and depth map is exploited, the compression efficiency may be improved compared with encoding them independently using H.264/AVC. A new encoder algorithm which employs Scalable Video Coding (SVC), the scalable extension of H.264/AVC, to compress the texture video and its associated depth map is proposed in this paper. Experimental results show that the proposed algorithm can provide up to 0.97 dB gain for the coded depth maps, compared with the simulcast scheme, wherein texture video and depth map are coded independently by H.264/AVC. Siping Tao, Ying Chen 0011, Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj, Houqiang Li |
ISCAS | 5 |
| 2009 | Media aware FEC for Ccalable Video Coding transmissionabstractScalable video coding (SVC) is an extension to H.264/AVC that enables encoding a video sequence once and using it subsequently at multiple operations points by different devices and applications. Forward Error Correction techniques may then be applied to the encoded video in order to enhance the robustness against transmission errors. Media aware FEC code construction caters for scalable video streams by adjusting the protection level for each video layer. In this paper, we discuss different approaches for media aware Forward Error Correction using Scalable Video Coded video that is transmitted over error prone channels. We describe two recent solutions for Unequal Error Protection and then evaluate those through multiple simulations. Lukasz Kondrad, Imed Bouazizi, Moncef Gabbouj |
ISCC | 3 |
| 2009 | Perceptual quality assessment based on visual attention analysisabstractMost existing quality metrics do not take the human attention analysis into account. Attention to particular objects or regions is an important attribute of human vision and perception system in measuring perceived image and video qualities. This paper presents an approach for extracting visual attention regions based on a combination of a bottom-up saliency model and semantic image analysis. The use of PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural SIMilarity) in extracted attention regions is analyzed for image/video quality assessment, and a novel quality metric is proposed which can exploit the attributes of visual attention information adequately. The experimental results with respect to the subjective measurement demonstrate that the proposed metric outperforms the current methods. Junyong You, Andrew Perkis, Miska M. Hannuksela, Moncef Gabbouj |
ACM Multimedia | 4 |
| 2009 | Coding techniques in Multiview Video Coding and Joint Multiview Video ModelabstractSince early 2006, Joint Video Team has been devoting on the development of Multiview Video Coding (MVC) standard as an extension of H.264/AVC. This MVC standard has been finalized in 2008. During the standardization of MVC, there was also a project namely Joint Multiview Video Model (JMVM), which focused on the advanced coding tools that are potentially useful. Those coding tools adopted into JMVM, including illumination compensation and motion skip, have not been added into MVC specification. In this paper, coding techniques in MVC as well as the tools in JMVM are described and discussed, focusing on the coding efficiency. Ying Chen 0011, Miska M. Hannuksela, Antti Hallapuro, Moncef Gabbouj, Houqiang Li |
PCS | 5 |
| 2009 | Efficient hierarchical inter picture coding for H.264/AVC baseline profileabstractBi-predictive (B) slices are not supported in the Baseline profile of the Advanced Video Coding (H.264/AVC) standard, which results in a decreased coding efficiency compared with other profiles supporting B slices. However, many application standards, such as the mobile multimedia services specified by the Third Generation Partnership Project (3GPP), use only the Baseline profile for H.264/AVC. Therefore, it is worth investigating H.264/AVC coding when only intra (I) and inter (P) slices are supported. In this paper, a content-adaptive Quantization Parameter (QP) cascading scheme for the hierarchical P coding method compatible with Baseline profile of H.264/AVC is proposed. The proposed method is based on a picture-level QP optimization. The proposed method has a significantly better rate-distortion performance than the traditional IPPP coding structure and outperforms hierarchical P coding methods using fixed delta QP settings between temporal levels noticeably with up to 0.53 dB gain in average luminance Peak Signal-to-Noise Ratio (PSNR). Weixing Wan, Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj |
PCS | 6 |
| 2009 | Low complexity algorithm for Spatially Varying TransformsabstractIn our previous work, we introduced spatially varying transforms (SVT) for video coding, where the location of the transform block within the macroblock is not fixed but varying. SVT has lower decoding complexity compared to standard methods as only a portion of the prediction error needs to be decoded. However, the encoding complexity of SVT can be relatively high because of the need to perform rate distortion optimization (RDO) for each candidate location parameter (LP). In this work, we propose a low complexity algorithm operating on macroblock and block level to reduce the encoding complexity of SVT. The proposed low complexity algorithm includes selection of available candidate LP based on motion difference and a hierarchical search algorithm. Experimental results show that the proposed low complexity algorithm can reduce around 80% of the candidate LP tested in RDO with only marginal penalty in coding efficiency. Cixun Zhang, Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
PCS | 5 |
| 2009 | Video Coding Using Spatially Varying Transform
Cixun Zhang, Kemal Ugur, Jani Lainema, Moncef Gabbouj |
PSIVT | 4 |
| 2009 | Mobile and Interactive Social Television - A virtual TV roomabstractSmart phones are becoming more and more powerful. Services that were traditionally designed for a static environment can now be implemented into mobile devices. One such service is Interactive Social TV, which allows geographically dispersed people to meet in a virtual shared space and watch TV while being able to interact with each other. This paper presents two novel architectures of a Mobile and Interactive Social TV system. In both of the architectures, the interaction is represented by a rich audio-visual media, allowing users to hear and see each other. In the first architecture, the mixing of the TV content with the interaction media is performed at the server side. In the second architecture, the mixing is performed in each client device. The issues of decoding and rendering simultaneous content and interaction media streams on a mobile device are discussed, and the related implementation is presented. Francesco Cricri, Sujeet Mate, Igor D. D. Curcio, Moncef Gabbouj |
WOWMOM | 4 |
| 2009 | Evolutionary artificial neural networks by multi-dimensional particle swarm optimization
Serkan Kiranyaz, Turker Ince, Alper Yildirim, Moncef Gabbouj |
Neural Networks | 4 |
| 2009 | Error Resilient Coding and Error Concealment in Scalable Video CodingabstractScalable video coding (SVC), which is the scalable extension of the H.264/AVC standard, was developed by the Joint Video Team (JVT) of ISO/IEC MPEG (Moving Picture Experts Group) and ITU-T VCEG (Video Coding Experts Group). SVC is designed to provide adaptation capability for heterogeneous network structures and different receiving devices with the help of temporal, spatial, and quality scalabilities. It is challenging to achieve graceful quality degradation in an error-prone environment, since channel errors can drastically deteriorate the quality of the video. Error resilient coding and error concealment techniques have been introduced into SVC to reduce the quality degradation impact of transmission errors. Some of the techniques are inherited from or applicable also to H.264/AVC, while some of them take advantage of the SVC coding structure and coding tools. In this paper, the error resilient coding and error concealment tools in SVC are first reviewed. Then, several important tools such as loss-aware rate-distortion optimized macroblock mode decision algorithm and error concealment methods in SVC are discussed and experimental results are provided to show the benefits from them. The results demonstrate that PSNR gains can be achieved for the conventional inter prediction (IPPP) coding structure or the hierarchical bi-predictive (B) picture coding structure with large group of pictures size, for all the tested sequences and under various combinations of packet loss rates, compared with the basic joint scalable video model (JSVM) design applying no error resilient tools at the encoder and only picture copy error concealment method at the decoder. Ying Chen 0011, Ye-Kui Wang, Houqiang Li, Miska M. Hannuksela, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2009 | Video Coding With Low-Complexity Directional Adaptive Interpolation FiltersabstractA novel adaptive interpolation filter structure for video coding with motion-compensated prediction is presented in this letter. The proposed scheme uses an independent directional adaptive interpolation filter for each sub-pixel location. The Wiener interpolation filter coefficients are computed analytically for each inter-coded frame at the encoder side and transmitted to the decoder. Experimental results show that the proposed method achieves up to 1.1 dB coding gain and a 15% average bit-rate reduction for high-resolution video materials compared to the standard nonadaptive interpolation scheme of H.264/AVC, while requiring 36% fewer arithmetic operations for interpolation. The proposed interpolation can be implemented in exactly 16-bit arithmetic, thus it can have important use-cases in mobile multimedia environments where the computational resources are severely constrained. Dmytro Rusanovskyy, Kemal Ugur, Antti Hallapuro, Jani Lainema, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2008 | LSF mapping for voice conversion with very small training setsabstractTo make voice conversion usable in practical applications, the number of training sentences should be minimized. With traditional Gaussian mixture model (GMM) based techniques small training sets lead to over-fitting and estimation problems. We propose a new approach for mapping line spectral frequencies (LSFs) representing the vocal tract. The idea is based on inherent intra-frame correlations of LSFs. For each target LSF, a separate GMM is used and only the source and target LSF elements best correlating with the current LSF are used in training. The proposed method is evaluated both objectively and in listening tests, and it is shown that the method outperforms the conventional GMM approach especially with very small training sets. Elina Helander, Jani Nurminen, Moncef Gabbouj |
ICASSP | 3 |
| 2008 | A detection algorithm for zero-quantized DCT coefficients in JPEGabstractThe discrete cosine transform (DCT) is widely used in image/video coding standards. However, since most DCT coefficients will be quantized to zeros, a large number of redundant computations are introduced. This paper presents an early detection algorithm to predict zero-quantized DCT coefficients for fast JPEG encoding. Based on the theoretical analysis for 2-D DCT and quantization in JPEG standard, we derive a sufficient condition under which each quantized coefficient becomes zero. Finally, the transform of the zero-quantized coefficients is omitted. Experimental results show that the proposed algorithm can significantly reduce the redundant computations and speed up the image encoding. Moreover, it doesn’t cause any performance degradation. Computational reduction also implies longer battery lifetime and energy economy for digital applications. Jin Li 0006, Jarmo Takala, Moncef Gabbouj, Hexin Chen |
ICASSP | 3 |
| 2008 | Efficient calculation of adaptive interpolation filter with distortion modellingabstractA novel method is proposed to calculate the coefficients of adaptive interpolation filter used in hybrid video coders for improving the coding efficiency. The proposed algorithm first selects the motion blocks where the majority of prediction errors result from mismatches in motion estimation and from aliasing present in the signal. This is realized by using a second order distortion model to estimate the effect of quantization on motion prediction error and coding results of the previous frames. Then, the filter coefficients are calculated analytically by minimizing the prediction error of those selected blocks. Experimental results show that the proposed method achieves up-to 0.6 dB gain compared to the standard H.264/AVC. Compared to other methods that calculate the filter coefficients using all motion blocks of the frame, the proposed method has significantly less encoding complexity (83% on average) with practically no penalty on coding efficiency. Kemal Ugur, Dmytro Rusanovskyy, Moncef Gabbouj |
ICASSP | 3 |
| 2008 | Video coding using pruned transforms and interleaving of multiple blockabstractTechnologies used in today's video coding standards have been designed and optimized mainly for standard definition (SD) resolutions and below. When moving to higher resolutions, data in video frames tend to become more correlated spatially. In this paper, we study how to take advantage of this phenomenon to lower computational requirements for high definition (HD) video coding. A coding method based on low complexity pruned transforms and interleaving of multiple transform coefficient blocks is proposed. An example implementation of this method in the context of H.264/AVC is also presented. Experimental results show that the proposed method maintains the high compression efficiency of H.264/AVC while significantly lowering the coding complexity of the codec. Cixun Zhang, Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
ICASSP | 5 |
| 2008 | Picture-level adaptive filter for asymmetric stereoscopic videoabstractIn asymmetric stereoscopic video coding, one view is coded in a quarter of the resolution of the other and the low- resolution view is predicted from the high-resolution view. This way, stereoscopic video effect could be achieved with only moderately increased bandwidth and complexity. Inter-view prediction tools for generating the predictor of a maroblock (MB) or MB partition in the low-resolution view from the high-resolution view play a vital role for coding efficiency in asymmetric video coding. In this paper, we propose a method that applies an adaptive filter to generate picture-level adaptive inter-view predictors for MBs or MB partitions. At the encoder, a low complexity preprocessing module is built to find out the filters. Simulation results show that the proposed method provides a bit-rate saving of 26% at maximum and 5% on average. Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj |
ICIP | 4 |
| 2008 | Low-complexity asymmetric multiview video codingabstractMultiview video coding (MVC) is currently under development by the Joint Video Team (JVT) as an extension to Advanced Video Coding (H264/AVC). Based on the suppression theory in binocular vision, the fidelity of one of the two views of a stereoscopic display can be reduced without noticeable degradation of subjective quality. Thus, in MVC, a subset of views can be coded with lower spatial resolution at negligible cost to subjective quality. Due to different resolutions, a downsampling process is required in an MVC decoder in order to enable motion compensation (MC) between views. In this paper, a low-complexity MC algorithm is proposed for MVC to enable inter-view prediction between pictures with different resolutions. It requires lower memory consumption and lower computational complexity compared with the conventional downsampled inter-view prediction, while providing comparable efficiency, as shown by the simulation results. Ying Chen 0011, Shujie Liu 0001, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj |
ICME | 6 |
| 2008 | Single-loop decoding for multiview video codingabstractMultiview video coding (MVC) is currently being standardized by the Joint Video Team as an extension of H264/AVC. When an MVC bitstream is decoded, some views (named target views) are to be displayed; some other views (named dependent views) may not be displayed but are needed for inter-view prediction of the target views. The original MVC design requires pictures of the dependent views to be fully decoded and stored. This entails both high decoding complexity and high memory consumption for the pictures in the views which are not intended for display, particularly when the number of dependent views is large. In this paper, a single-loop decoding (SLD) scheme is introduced to address these disadvantages. SLD requires only partial decoding of pictures in dependent views and thus significantly reduces decoding complexity and memory consumption. The proposed method is based on the so-called motion skip, wherein inter-view motion and coding mode prediction is exploited. Experimental results show that compared to coding schemes that require comparable complexity, significant compression gain can be achieved. For example, 25% bit-rate saving on average can be obtained compared to simulcast. Simulation results also show that the proposed SLD scheme provides a substantial reduction of complexity and memory size, at the expense of only a minor compression efficiency loss, compared with multiple-loop decoding MVC schemes. Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj |
ICME | 4 |
| 2008 | Laplacian modeling of DCT coefficients for real-time encodingabstractDigital image/video coding standards such as JPEG, H.264 are becoming more and more important for multimedia applications. Due to the huge amount of computations, there are significant efforts to speed up the encoding process. This paper proposes a Laplacian based statistical model to predict zero-quantized DCT coefficients in JPEG and to reduce the computations of encoding process. Compared with the standard JPEG and the reference in the literature, the proposed model can significantly simplify the computational complexity and achieve the best real-time performance at the expense of negligible visual degradation. Moreover, it can be directly applied to other DCT-based image/video codec. Computational reduction also implies longer battery lifetime and energy economy for digital applications. Jin Li 0006, Moncef Gabbouj, Jarmo Takala, Hexin Chen |
ICME | 2 |
| 2008 | Fast mode decision for adaptive prediction error codingabstractIn [1], Adaptive Prediction Error Coding (APEC) in spatial and frequency domain is proposed and significantly improves the coding efficiency of video coders. However, this approach comes with the expense of increased encoding complexity because of the need to perform Rate Distortion Optimization (RDO) for each block, to decide whether the block is coded in spatial or frequency domain. In this work, we propose novel fast mode decision algorithm operating on macroblock and block level to reduce the encoding complexity of APEC. The proposed algorithm estimates the edge orientation of the residual block using spatial domain filtering, and based on the estimated edge orientation, reduces the number of candidates to be checked in RDO. Furthermore, proposed method includes an early termination step that stops searching the best candidate based on the difference between spatial and frequency domain coding. Experimental results show that the proposed fast mode decision algorithm can achieve averagely 90.7% encoding time reduction of APEC with only about 0.06 dB loss in coding efficiency. Cixun Zhang, Kemal Ugur, Moncef Gabbouj |
ICME | 3 |
| 2008 | Unsupervised design of Artificial Neural Networks via multi-dimensional Particle Swarm OptimizationabstractIn this paper, we present a novel and efficient approach for automatic design of artificial neural networks (ANNs) by evolving to the optimal network configuration(s) within an architecture space. The evolution technique, the so-called multidimensional particle swarm optimization (MD PSO) re-forms the native structure of PSO particles in such a way that they can make inter-dimensional passes with a dedicated dimensional PSO process. So in a multidimensional search space where the optimum dimension is unknown, swarm particles can seek for both positional and dimensional optima. This eventually removes the necessity of setting a fixed dimension a priori, which is a common drawback for the family of swarm optimizers. With the proper encoding of the network configurations and parameters into particles, MD PSO can then seek for positional optimum in the error space and dimensional optimum in the architecture space. The optimum dimension converged at the end of a MD PSO process corresponds to a unique ANN configuration where the network parameters (connections, weights and biases) can then be resolved from the positional optimum reached on that dimension. The efficiency and performance of the proposed technique is demonstrated over one of the hardest synthetic problems. The experimental results show that MD PSO evolves to optimum or near-optimum networks in general. Serkan Kiranyaz, Turker Ince, Alper Yildirim, Moncef Gabbouj |
ICPR | 4 |
| 2008 | On the impact of alignment on voice conversion performanceabstractMost of the current voice conversion systems model the joint density of source and target features using a Gaussian mixture model. An inherent property of this approach is that the source and target features have to be properly aligned for the training. It is intuitively clear that the accuracy of the alignment has some effect on the conversion quality but this issue has not been thoroughly studied in the literature. Examples of alignment techniques include the usage of a speech recognizer with forced alignment or dynamic time warping (DTW). In this paper, we study the effect of alignment on voice conversion quality through extensive experiments and discuss issues that should be considered. The main outcome of the study is that alignment clearly matters but with simple voice activity detection, DTW and some constraints we can achieve the same quality as with hand-marked labels. Elina Helander, Jan Schwarz, Jani Nurminen, Hanna Silén, Moncef Gabbouj |
INTERSPEECH | 5 |
| 2008 | Evaluation of Finnish unit selection and HMM-based speech synthesis
Hanna Silén, Elina Helander, Jani Nurminen, Moncef Gabbouj |
INTERSPEECH | 4 |
| 2008 | Frame loss error concealment for multiview video codingabstractThe Multiview Video Coding (MVC) standard is currently under development by the Joint Video Team as an extension of the Advanced Video Coding (H.264/AVC) standard. An MVC encoder compresses more than one viewpoint of a scene captured by different cameras. Redundancies between views can be used for inter-view prediction in encoding as well as error concealment in decoding. In this paper, a new algorithm utilizing motion information of pictures from other views to conceal a lost picture is proposed. The algorithm first derives motion information for a lost picture based on motion fields of pictures in adjacent views. Then, traditional motion compensation is invoked within the view containing the lost picture to derive a concealed frame. Experimental results show that the proposed algorithm can improve video quality with a negligible computational complexity overhead compared to simple temporal error concealment algorithms. Shujie Liu 0001, Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela, Houqiang Li |
ISCAS | 4 |
| 2008 | Video coding with pixel-aligned directional adaptive interpolation filtersabstractIn this paper a novel adaptive interpolation filter structure is proposed to improve the coding efficiency of video coders. Proposed scheme utilizes one dimensional directional adaptive filter for every of the sub-pixel location, whose coefficients are calculated analytically for every frame by minimizing the prediction error energy. The direction of the interpolation filter is different for every sub-pixel position and it is determined based on the alignment of the corresponding sub- pixel with integer pixel samples. Experimental results show that, the proposed method achieves up-to 1.1 dB gain compared to the standard non-adaptive interpolation scheme of H.264/AVC, requiring less number of operations for interpolation. Compared to two-dimensional non-separable adaptive interpolation, proposed scheme has practically the same coding efficiency with approximately 3 times less complexity. Since significant coding efficiency is achieved without increasing the complexity, it is believed that proposed method has important use-cases in mobile multimedia environments where the resources are severely constrained. Dmytro Rusanovskyy, Kemal Ugur, Moncef Gabbouj, Jani Lainema |
ISCAS | 3 |
| 2008 | Fast encoding algorithms for video coding with adaptive interpolation filtersabstractIn order to compensate for the temporally changing effect of aliasing and improve the coding efficiency of video coders, adaptive interpolation filtering schemes have been recently proposed. In such schemes, encoder computes the interpolation filter coefficients for each frame and then re-encodes the frame with the new adaptive filter. However, the coding efficiency benefit comes with the expense of increased encoding complexity due to this additional encoding pass. In this paper, we present two novel algorithms to reduce the encoding complexity of adaptive interpolation filtering schemes. First algorithm reduces the complexity of the second encoding pass by using a very lightweight motion estimation algorithm that reuses the data already computed in the first encoding pass. Second algorithm eliminates the second coding pass and re-uses the filter coefficients already computed for previous frames. Experimental results show that the proposed methods achieve between 1.5 to 2 times encoding complexity reduction with practically negligible penalty on coding efficiency. Dmytro Rusanovskyy, Kemal Ugur, Moncef Gabbouj |
MMSP | 3 |
| 2008 | Seamless handover for mobile TV over DVB-H applicationsabstractIn this paper, we present an approach for achieving application driven seamless handover across different wireless access networks. The target is thereby to eliminate any content loss while minimizing the interruption duration that is perceived by the user during the handover procedure. Ultimately, the user experience will greatly benefit from the usage of the most suitable network access technology. The focus is set on Mobile TV applications, where an ongoing streaming session experiences an application-controlled handover from DVB-H to an alternative unicast access (e.g. Wi-Fi or 3G) or vice versa. A handover decision algorithm is also presented, aiming at enhancing the overall system stability by adjusting to the dynamic conditions of the broadcast channel. Lukasz Kondrad, Imed Bouazizi, Moncef Gabbouj |
MUM | 3 |
| 2008 | Semi-Fuzzy Rate Controller for Variable Bit Rate VideoabstractA novel semi-fuzzy (SF) rate control algorithm (RCA) for variable bit rate (VBR) video applications is proposed. The proposed RCA is optimized to provide high quality compressed video bit streams in a wide operating range from constant quality to nearly constant bit rate. Thanks to a low degree of computational complexity, it is suitable for real-time applications of VBR video. The proposed RCA operates under given buffer size, delay and quality constraints. It provides a VBR video bit stream by controlling the quantization parameter (QP) on a picture basis. The QP is mainly controlled by a fuzzy rate controller and a deterministic quality controller, which are optimized such that they minimize the variation of quality to provide encoded video with high and stable visual quality. The proposed RCA has been implemented in an H.264/AVC video codec and the experimental results show that it provides a high-level average quality for encoded video while strictly obeying the buffering delay and quality constraints. Mehdi Rezaei, Miska M. Hannuksela, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | A Generic Shape/Texture Descriptor Over Multiscale Edge Field: 2-D Walking Ant HistogramabstractA novel shape descriptor, which can be extracted from the major object edges automatically and used for the multimedia content-based retrieval in multimedia databases, is presented. By adopting a multiscale approach over the edge field where the scale represents the amount of simplification, the most relevant edge segments, referred to as subsegments, which eventually represent the major object boundaries, are extracted from a scale-map. Similar to the process of a walking ant with a limited line of sight over the boundary of a particular object, we traverse through each subsegment and describe a certain line of sight, whether it is a continuous branch or a corner, using individual 2-D histograms. Furthermore, the proposed method can also be tuned to be an efficient texture descriptor, which achieves a superior performance especially for directional textures. Finally, integrating the whole process as feature extraction module into MUVIS framework allows us to test the mutual performance of the proposed shape descriptor in the context of multimedia indexing and retrieval. Serkan Kiranyaz, Miguel Ferreira, Moncef Gabbouj |
IEEE Trans. Image Process. | 3 |
| 2008 | Joint Video Coding and Statistical Multiplexing for Broadcasting Over DVB-H ChannelsabstractA novel joint video encoding and statistical multiplexing (StatMux) method for broadcasting over digital video broadcasting for handhelds (DVB-H) channels is proposed to improve the quality of encoded video and to decrease the end-to-end delay in a broadcast system. The main parts of end-to-end delay in a DVB-H system result from a time-sliced transmission scheme that is used in DVB-H and from the bit rate variations of service bit streams. The time-sliced transmission scheme is utilized in DVB-H to reduce the power consumption of DVB-H receivers. Variable bit rate (VBR) video bit streams are used in DVB-H to improve the video quality and compression performance. The time-sliced transmission scheme has increased the channel switching delay, i.e., switching to a new audio-visual service, in DVB-H. The used VBR bit streams increase the required buffering delays in the whole system. The different parts of end-to-end delay in a DVB-H system can be affected by the used video encoding and multiplexing methods. Different scenarios for encoding and StatMux of video sources for DVB-H application are studied in this paper. Moreover, a new method for jointly encoding and StatMux of video sources is proposed that not only decreases the end-to-end delay but also improves the average quality of compressed video by dynamically distributing available bandwidth between the video sources according to their relative complexity. Performance of the proposed method is validated by simulation results. Mehdi Rezaei, Imed Bouazizi, Moncef Gabbouj |
IEEE Trans. Multim. | 3 |
| 2007 | Adaptive Interpolation Filter with Flexible Symmetry for Coding High Resolution High Quality VideoabstractIn this work, a novel sub-pixel interpolation algorithm is proposed for video coders targeted towards high resolution and high fidelity use cases. Proposed scheme is based on adapting the interpolation filter's symmetry assumptions in a rate-distortion-optimized fashion, taking into account the coding rate and the statistical properties of each image of the video sequence. Experimental results show that, using the proposed algorithm a gain of up-to 1.1 dB is achieved compared to non-adaptive sub-pixel interpolation of H.264/AVC. Compared to other state-of-the-art adaptive sub-pixel interpolation methods, a gain of up-to 0.5 dB is achieved. Proposed scheme outperforms H.264/AVC for all test cases; however, improvement is more significant at high bitrates and at high resolutions. This is especially important for future video coding solutions targeting high fidelity video applications. Kemal Ugur, Jani Lainema, Moncef Gabbouj |
ICASSP (1) | 3 |
| 2007 | Spatial and Temporal Adaptation of Interpolation Filter For Low Complexity Encoding/DecodingabstractCompared to video coding with non-adaptive interpolation filtering, adaptive filters achieve higher compression ratios, with an increase in encoding and decoding complexity. In our earlier work, we significantly reduced the decoding complexities of adaptive filtering schemes with a minimal impact on the coding efficiency by making use of different filters and adapting them spatially and temporally. However, our previous scheme required high encoder complexity, as several encoding passes per frame were needed to analyze the input image and optimize the selection of interpolation filters. In this paper, a novel algorithm that does not require multiple encoding passes, but still give similar or better performance is proposed. This is achieved by using a modified decision making function that does not require full reconstruction of coded frame and use motion and prediction information more efficiently. In addition, we generalized our previous scheme by introducing additional filters, so that better Rate-Distortion-Complexity tradeoffs are possible. Experimental results show that up-to 50-70% reduction in interpolation complexity is achieved, with less than 0.13 dB penalty on coding efficiency. Dmytro Rusanovskyy, Moncef Gabbouj, Kemal Ugur |
MMSP | 2 |
| 2007 | Hierarchical Cellular Tree: An Efficient Indexing Scheme for Content-Based Retrieval on Multimedia DatabasesabstractOne of the challenges in the development of a content-based multimedia indexing and retrieval application is to achieve an efficient indexing scheme. The developers and users who are accustomed to making queries to retrieve a particular multimedia item from a large scale database can be frustrated by the long query times. Conventional indexing structures cannot usually cope with the requirements of a multimedia database, such as dynamic indexing or the presence of high-dimensional audiovisual features. Such structures do not scale well with the ever increasing size of multimedia databases whilst inducing corruption and resulting in an over-crowded indexing structure. This paper addresses such problems and presents a novel indexing technique, hierarchical cellular tree (HCT), which is designed to bring an effective solution especially for indexing large multimedia databases. Furthermore it provides an enhanced browsing capability, which enables user to make a guided tour within the database. A pre-emptive cell-search mechanism is introduced in order to prevent corruption, which may occur due to erroneous item insertions. Among the hierarchical levels that are built in a bottom-up fashion, similar items are collected into appropriate cellular structures at some level. Cells are subject to mitosis operations when the dissimilarity exceeds a required level. By mitosis operations, cells are kept focused and compact and yet, they can grow into any dimension as long as the compactness is maintained. The proposed indexing scheme is then used along with a recently introduced query method, the progressive query, in order to achieve the ultimate goal, from the user point of view that is retrieval of the most relevant items in the earliest possible time regardless of the database size. Experimental results show that the speed of retrievals is significantly improved and the indexing structure shows no sign of degradations when the database size is increased. Furthermore, HCT indexing body can conveniently be used for efficient browsing and navigation operations among the multimedia database items Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Multim. | 2 |
| 2006 | Multi-Scale Edge Detection and Object Extraction for Image RetrievalabstractA new scheme for boundary-based object extraction and description of still images, through multi-scale edge detection, is proposed in this paper. Boundary-based methods try to extract closed contours from individual edge pixels through edge-linking. Our approach is based on a connected structure of edge pixels as the initial edge-linking elements. These connected structures, the sub-segments, are extracted from the Canny edge map of an image. Multiple simplification-scales are derived from applying iterations of the bilateral filter to the image, providing extra information about the relative importance of each sub-segment. Edge-linking towards contour closure is achieved through perceptually-driven minimum cost search. Furthermore, a shape-based description vector is derived from the extracted contours, and retrieval results are obtained via the integration of the whole scheme into MUVIS framework Miguel Ferreira, Serkan Kiranyaz, Moncef Gabbouj |
ICASSP (2) | 3 |
| 2006 | Low-Complexity Fuzzy Video Rate Controller for StreamingabstractIn this paper we propose a low-complexity fuzzy video rate control algorithm with buffer constraint designed for real-time streaming applications. While in low delay video communications bit streams with constant bitrate are required, in streaming application more delay and variation in bitrate is acceptable. The described video rate control algorithm (RCA) provides a variable bitrate video by control of the quantization scale (QS) on picture basis. The QS is mainly controlled by a fuzzy controller such that it minimizes the variation of QS to provide encoded video with high visual quality so as to utilize the variable bitrate benefits as much as possible. The proposed rate control algorithm (RCA) has been implemented in the MPEG-4, H.263 and H.264/AVC standard video codecs and the experimental results show that it provides high level average quality for encoded video while it strictly obeys streaming constraints Mehdi Rezaei, Miska M. Hannuksela, Moncef Gabbouj |
ICASSP (2) | 3 |
| 2006 | Generating H.264/AVC Compliant Bitstreams for Lightweight Decoding Operation Suitable for Mobile Multimedia SystemsabstractIn this work, we propose novel encoder algorithms for the state-of-the-art video coding standard H.264, to generate decoder friendly video bitstreams. Using the proposed algorithms, it is possible to generate bitstreams requiring significantly less decoding complexity, with negligible effect on picture quality. This is achieved by using novel algorithms for mode decision and motion estimation that bias easy-to-decode motion vectors in a rate-distortion optimized fashion. Experimental results show that, more than 15% decoding complexity reduction is achieved with less than a 0.1 dB penalty on the average video quality. We believe that this approach has potential in various use cases especially in mobile multimedia systems, where the video decoder operation is often dominating the handsets power consumption Kemal Ugur, Jani Lainema, Antti Hallapuro, Moncef Gabbouj |
ICASSP (2) | 4 |
| 2006 | Fuzzy Rate Controller for Variable Bitrate Video in Mobile ApplicationsabstractIn this paper we propose a low-complexity fuzzy video rate controller designed for real-time variable bitrate applications with buffer constraints. The algorithm is optimized for streaming application in mobile devices. Furthermore, the proposed algorithm can be used for local recording application so as the recorded video may be streamed in future. Today, many mobile phones include a digital camera that can be used to capture and encode video in real-time. We assume that no memory for storage of uncompressed video is available in mobile phones. Therefore, look-ahead and multi-pass rate controls are not possible. Furthermore, considering the processing power and, more importantly, battery life constraints in mobile devices, the proposed algorithm needs to be as simple as possible. The described variable bitrate (VBR) bit rate control algorithm controls the quantization scale (QS) on picture basis. The QS is mainly controlled by a fuzzy controller such that it minimizes the variation of QS to provide encoded video with high visual quality so as to utilize the variable bitrate benefits as much as possible. The proposed rate control algorithm (RCA) has been implemented in the MPEG-4, H.263 and H.264/AVC standard video codecs and the experimental results show that it provides high level average quality for encoded video while it strictly obeys buffering constraints. Mehdi Rezaei, Alireza Akhbardeh, Miska M. Hannuksela, Moncef Gabbouj |
ICC | 4 |
| 2006 | Video Splicing and Fuzzy Rate Control in IP Multi-Protocol Encapsulator for Tune-In Time Reduction in IP Datacasting (IPDC) over DVB-HabstractA novel video splicing and rate control method is proposed which minimizes the tune-in time in IPDC over DVB-H. DVB-H uses a time-sliced transmission scheme to reduce the power consumption used for radio reception. One of the significant factors in tune-in time is the time from the start of media decoding to the start of correct output from decoding, which would be minimized when a time-slice is started with a random access point picture such as an instantaneous decoding refresh (IDR) picture in H.264/AVC. In IPDC over DVB-H, the encapsulation to time-slices is performed independently of encoding in a network element called IP encapsulator. At the time of encoding, time-slice boundaries are not known exactly, and it is impossible to govern the location of IDR pictures relative to time-slices. It is proposed that an additional stream consisting of IDR pictures only is transmitted to the IP encapsulator, which replaces pictures in a normal bitstream with IDR pictures according to time-slice boundaries in order to achieve the minimum tune-in time. It has to be ensured that the "spliced" bit stream resulting from the operation of the IP encapsulator complies with the hypothetical reference decoder (HRD) specification of H.264/AVC. A video rate control system utilizing a fuzzy controller is proposed to satisfy the HRD requirements for the spliced bit stream. Simulation results show that the proposed splicing method and rate control system can provide standard bit streams with good average quality of decoded video and with minimum tune-in time. Mehdi Rezaei, Miska M. Hannuksela, Moncef Gabbouj |
ICIP | 3 |
| 2006 | Video Encoding and Splicing for Tune-in Time Reduction in IP Datacasting (IPDC) Over DVB-HabstractA novel video encoding and splicing method is proposed which minimizes the tune-in time of "channel zapping", i.e. changing from one audiovisual service to another, in IPDC over digital video broadcasting for handheld terminals (DVB-H). DVB-H uses a time-sliced transmission scheme to reduce the power consumption used for radio reception. Tune-in time in DVB-H refers to the time between the start of the reception of a broadcast signal and the start of the media rendering. One of the significant factors in tune-in time is the time from the start of media decoding to the start of correct output from decoding, which is minimized when a time-slice is started with a random access point picture such as an independent decoding refresh (IDR) picture in H.264/AVC. In IPDC over DVB-H, encapsulation to time-slices is performed independently from encoding in a network element called IP encapsulator. At the time of encoding, time-slice boundaries are not known exactly, and it is therefore impossible to govern the location of IDR pictures relative to time-slices. It is proposed that an additional stream consisting of IDR pictures only is transmitted to the IP encapsulator, which replaces pictures in a normal bitstream with IDR pictures according to time-slice boundaries in order to achieve the minimum tune-in time. It has to be ensured that the "spliced" stream resulting from the operation of the IP encapsulator complies with the hypothetical reference decoder (HRD) specification of H.264/AVC. A video encoding and rate control system is proposed to satisfy the HRD requirements for the spliced stream. Simulation results show that in addition to fulfilling HRD compliancy, good average quality of decoded video is achieved with minimum tune-in time Mehdi Rezaei, Miska M. Hannuksela, Moncef Gabbouj |
ICME | 3 |
| 2006 | Spliced Video and Buffering Considerations for Tune-In Time Minimization in DVB-H for Mobile TVabstractA novel video splicing method is proposed which minimizes the tune-in time of mobile TV in Digital Video Broadcasting for Handheld terminals (DVB-H). DVB-H uses a time-sliced transmission scheme to reduce the power consumption used for radio reception. Tune-in time in DVB-H refers to the time between the start of the reception of a broadcast signal and the start of the media rendering. One of the significant factors in tune-in time is the time from the start of media decoding to the start of correct output from decoder, which can be minimized when a time-slice is started with a random access point picture such as an independent decoding refresh (IDR) picture in H.264/AVC. In IP datacasting (IPDC) over DVB-H, the encapsulation to time-slices is performed independently from encoding in a network element called IP encapsulator. At the time encoding, time-slice boundaries are not known exactly, and it is impossible to govern the location of IDR pictures relative to time-slice boundaries. It is proposed that an additional stream consisting of IDR pictures only is transmitted to the IP encapsulator, which replaces pictures in a normal bitstream with IDR pictures according to time-slice boundaries in order to achieve the minimum tune-in time. It has to be ensured that the "spliced" stream resulting from the operation of the IP encapsulator complies with the Hypothetical Reference Decoder (HRD) specification of H.264/AVC. Mehdi Rezaei, Miska M. Hannuksela, Vinod Kumar Malamal Vadakital, Moncef Gabbouj |
PIMRC | 4 |
| 2006 | Method for Unequal Error Protection in DVB-H for Mobile TelevisionabstractThis paper introduces a method for unequal error protection (UEP) of media data in a time-sliced DVB-H channel. Media datagrams are assigned priorities using some a-priori knowledge. Datagrams covering a certain period of playback time are first grouped based on the priority assignment. Each group is then protected using Reed-Solomon forward error correction (FEC) codes and packed into multi-protocol encapsulation (MPE) FEC frames as defined by the DVB-H standard. All MPE-FEC frames for a certain period of playback time are then sent back to back without any delay between these MPE-FEC frames. This method of UEP is generic and can be tuned according to the priority assignment algorithm. Simulations using H.264/AVC video were conducted to evaluate the performance of the proposed method. It used a simple priority assignment algorithm. The resulting rate distortion graphs show good performance and an average luma peak signal-to-noise ratio (PSNR) improvement of up to 0.8 dB was achieved Vinod Kumar Malamal Vadakital, Miska M. Hannuksela, Mehdi Rezaei, Moncef Gabbouj |
PIMRC | 4 |
| 2006 | Multimedia indexing and retrieval: ever great challenges
Chaabane Djeraba, Moncef Gabbouj, Patrick Bouthemy |
Multim. Tools Appl. | 2 |
| 2006 | A generic audio classification and segmentation approach for multimedia indexing and retrievalabstractWe focus the attention on the area of generic and automatic audio classification and segmentation for audio-based multimedia indexing and retrieval applications. In particular, we present a fuzzy approach toward hierarchic audio classification and global segmentation framework based on automatic audio analysis providing robust, bi-modal, efficient and parameter invariant classification over global audio segments. The input audio is split into segments, which are classified as speech, music, fuzzy or silent. The proposed method minimizes critical errors of misclassification by fuzzy region modeling, thus increasing the efficiency of both pure and fuzzy classification. The experimental results show that the critical errors are minimized and the proposed framework significantly increases the efficiency and the accuracy of audio-based retrieval especially in large multimedia databases. Serkan Kiranyaz, Ahmad Farooq Qureshi, Moncef Gabbouj |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Automatic Object Extraction Over Multiscale Edge Field for Multimedia RetrievalabstractIn this work, we focus on automatic extraction of object boundaries from Canny edge field for the purpose of content-based indexing and retrieval over image and video databases. A multiscale approach is adopted where each successive scale provides further simplification of the image by removing more details, such as texture and noise, while keeping major edges. At each stage of the simplification, edges are extracted from the image and gathered in a scale-map, over which a perceptual subsegment analysis is performed in order to extract true object boundaries. The analysis is mainly motivated by Gestalt laws and our experimental results suggest a promising performance for main objects extraction, even for images with crowded textural edges and objects with color, texture, and illumination variations. Finally, integrating the whole process as feature extraction module into MUVIS framework allows us to test the mutual performance of the proposed object extraction method and subsequent shape description in the context of multimedia indexing and retrieval. A promising retrieval performance is achieved, and especially in some particular examples, the experimental results show that the proposed method presents such a retrieval performance that cannot be achieved by using other features such as color or texture. Serkan Kiranyaz, Miguel Ferreira, Moncef Gabbouj |
IEEE Trans. Image Process. | 3 |
| 2005 | Visual media retrieval using transform-based layered query schemeabstractThis paper presents a visual media querying scheme referred to as transform-based layered query (TLQ) scheme. The TLQ scheme mainly aims at decreasing retrieval processing time and run-time memory consumption without degrading retrieval results semantically. The scheme contains abstract layers in indexing and retrieval phases, where each indexing layer corresponds to a retrieval layer. The layers are constructed based on transformations for reducing visual frame and feature data dimensions. The proposed TLQ scheme also involves an unsupervised method for eliminating irrelevant media items between the retrieval layers. A two-layer TLQ system is implemented and integrated into MUVIS content-based multimedia indexing and retrieval framework, and its theoretical advantages are verified with dedicated experiments on image and video databases. The experiments reveal that 75% retrieval performance improvement in terms of process time can be achieved depending on transformation parameters. Esin Guldogan, Moncef Gabbouj, Olcay Guldogan |
ICIP (1) | 2 |
| 2005 | Block-based ordinal co-occurrence matrices for texture similarity evaluationabstractIn this paper we introduce a block-based approach for ordinal co-occurrence matrices aimed at improving robustness of the basic ordinal co-occurrence. Earlier, we have introduced two approaches for building ordinal co-occurrence matrices. One considers only the center pixel of a moving window as a seed point, compares it to its anti-causal neighbors and saves the occurrences of ordinal relations between pixels in the form of co-occurrence matrices. However, in that approach problems occur especially when considering textures with slightly varying gray levels in relatively large areas. The other method improves the robustness of the earlier method by considering also other pixels than the center pixel in the thresholded window as seed points. Main drawback of that method is the increased computational complexity. In order to avoid that, a block-based approach for building the ordinal co-occurrence matrices is introduced in this paper. Retrieval accuracy of the proposed method has shown to produce better results than existing techniques. Mari Partio, Bogdan Cramariuc, Moncef Gabbouj |
ICIP (1) | 3 |
| 2005 | An Efficient Image Retrieval Scheme on Java Enabled Mobile DevicesabstractContent-based image retrieval over wireless networks is a challenging research problem. In this paper, we present an efficient content-based image retrieval framework, which is developed for mobile platforms in client-server architecture and uses a combination of various low-level visual features. Several techniques were adapted in order to achieve the retrieval efficiency and query speed on mobile networks. Particularly a new implementation, which is called compact media retrieval on progressive query (CMR-PQ), is introduced. CMR-PQ is basically designed to retrieve the query results in an acceptable time, a network adaptive and configurable scheme. The experimental results present various query completion and image retrieval timings over several networks Iftikhar Ahmad 0001, Serkan Kiranyaz, Moncef Gabbouj |
MMSP | 3 |
| 2005 | Bit Allocation for Variable Bitrate VideoabstractIn this paper we propose a special bit allocation method which can be used in most rate control algorithms in variable rate video applications. In real-time video communication applications, we need a constant short-term average bitrate, while in variable bitrate applications such as streaming and local recording applications, a constant long-term average bitrate is sufficient and more short-term variation in bitrate is acceptable. In comparison with constant bitrate video, a variable bitrate video can provide better visual quality and coding efficiency for compressed video sequences. Furthermore, while more variation in bitrate is possible we have additional degrees of freedom to control the encoding parameters. We propose a special bit allocation algorithm to take advantage of this freedom in variable bitrate video. We introduce a new type of frame namely SPP frame (SPecial P frame) that can be used in combination with I, P, B and other types of frames in different encoders including H.263, MPEG-4 and H.264/AVC encoders. We propose a simple method to implement the SPP frames independently of the rate control algorithm. The experimental results show that the SPP frames can considerably increase the total average quality of variable rate encoded video Mehdi Rezaei, Moncef Gabbouj |
MMSP | 2 |
| 2005 | On Datacasting of H.264/AVC over DVB-HabstractThis paper investigates the performance of H. 264/AVC video codec in a DVB-H (Handheld) datacasting environment. DVB-H was designed to provide point-to-multipoint (PTM) broadcast/multicast type transmission to handheld, battery operated devices. Reed-Solomon (RS) forward error correcting (FEC) codes are applied to Multi-Protocol Encapsulation (MPE) section payloads, termed MPE-FEC, to deliver data over DVB-H. The bitrate overheads incurred due to the additional FEC and packetization headers are analyzed. The requirement of additional data protection in the form of MPE-FEC is illustrated with the help of simulation results. Vinod Kumar Malamal Vadakital, Miska M. Hannuksela, Harri Pekkonen, Moncef Gabbouj |
MMSP | 4 |
| 2005 | Video rate control for streaming and local recording optimized for mobile devicesabstractIn this paper, we propose a real-time, low-complexity video rate control algorithm designed to obey buffer constraints. The algorithm is optimized for streaming and local recording applications in mobile devices. Today, most mobile phones include a digital camera that can be used to capture video. The on-phone processor technology has become powerful enough to encode video in real-time. The resulting file can, for example, be archived in the phone's memory, or (progressively or as one block) downloaded, through the 3G mobile network, Bluetooth, or WLAN, to Internet-connected computer systems. From here, all forms of multimedia transmission, such as streaming, file sharing, or multimedia mail become possible. In local recording and streaming applications on a mobile phone, we assume that no memory for storage of uncompressed video is available. Therefore, look-ahead and multi-path rate control is not possible. Furthermore, considering the processing power and, more importantly, battery life constraints in mobile devices, the proposed algorithm needs to be as simple as possible. The described algorithm implements a variable bitrate (VBR) by controlling the quantization scale (QS) on a per picture basis. The QS is calculated based on two other QSs, which correspond to constant rate and constant quality rate controls. The algorithm utilizes the variable bitrate benefits as much as possible so as to minimize the variation of the QS scale, and to provide encoded video with high visual quality. Although it strictly obeys buffering constraints as discussed later, the experimental results show that it allows encoded video at average quality levels significantly higher than reported in earlier works Mehdi Rezaei, Stephan Wenger, Moncef Gabbouj |
PIMRC | 3 |
| 2004 | Modeling and real-time auralization of electrodynamic loudspeaker non-linearitiesabstractThe non-linear modeling of an electrodynamic speaker is studied and its use for real-time auralization of an arbitrary sound source is considered. First, the dominant observed non-linear behavior is presented. Then, a black-box approach to the system identification task is used seeking a generalized procedure applicable to any loudspeaker. In this respect, a separate identification scheme of the linear and non-linear characteristics of the loudspeaker is proposed. Some physical insight is used later in the model, which causes the loss of some of the desired generality. The model quality is finally assessed through subjective listening tests, and the results are presented. Marcelo Soria-Rodríguez, Moncef Gabbouj, Nick Zacharov, Matti S. Hämäläinen, Kalle Koivuniemi |
ICASSP (4) | 2 |
| 2004 | Texture similarity evaluation using ordinal co-occurrence
Mari Partio, Bogdan Cramariuc, Moncef Gabbouj |
ICIP | 3 |
| 2004 | Isolated regions in video codingabstractDifferent types of prediction are applied in modern video coding. While predictive coding improves compression efficiency, the propagation of transmission errors becomes more likely. In addition, predictive coding brings difficulties to other aspects of video coding, including random access, parallel processing, and scalability. In order to combat the negative effects, video coding schemes introduce mechanisms such as slices and intracoding, to limit and break the prediction. This paper proposes the use of the isolated regions coding tool that jointly limits in-picture prediction and interprediction on a region-of-interest basis. The tool can be used to provide random access points from non-intrapictures and to respond to intrapicture update requests. Furthermore, it can be applied as an error-robust macroblock mode decision method and can be used in combination with unequal error protection. Finally, it enables mixing of scenes, which is useful in coding of masked scene transitions. Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj |
IEEE Trans. Multim. | 3 |
| 2004 | Vector rational interpolation schemes for erroneous motion field estimation applied to MPEG-2 error concealmentabstractA study on the use of vector rational interpolation for the estimation of erroneously received motion fields of MPEG-2 predictively coded frames is undertaken in this paper, aiming further at error concealment (EC). Various rational interpolation schemes have been investigated, some of which are applied to different interpolation directions. One scheme additionally uses the boundary matching error and another one attempts to locate the direction of minimal/maximal change in the local motion field neighborhood. Another one further adopts bilinear interpolation principles, whereas a last one additionally exploits available coding mode information. The methods present temporal EC methods for predictively coded frames or frames for which motion information pre-exists in the video bitstream. Their main advantages are their capability to adapt their behavior with respect to neighboring motion information, by switching from linear to nonlinear behavior, and their real-time implementation capabilities, enabling them for real-time decoding applications. They are easily embedded in the decoder model to achieve concealment along with decoding and avoid post-processing delays. Their performance proves to be satisfactory for packet error rates up to 2% and for video sequences with different content and motion characteristics and surpass that of other state-of-the-art temporal concealment methods that also attempt to estimate unavailable motion information and perform concealment afterwards. Sofia Tsekeridou, Faouzi Alaya Cheikh, Moncef Gabbouj, Ioannis Pitas |
IEEE Trans. Multim. | 3 |
| 2003 | Relevance feedback for shape query refinementabstractIn this paper we propose to incorporate a feedback loop, into the ordinal correlation framework and apply it to shape-based image retrieval. The user's feedback on the relevance of the retrieval results is used to tune the weights of the similarity measure. Statistics from the features of both relevant and irrelevant items are used to estimate the weights. Moreover, the information accumulated from previous retrieval iterations is used in the weights estimation. A simple measure of the discrimination power is proposed and used to show that the relevance feedback increases the capability of the ordinal correlation scheme to discriminate between relevant and irrelevant objects. Faouzi Alaya Cheikh, Bogdan Cramariuc, Moncef Gabbouj |
ICIP (1) | 3 |
| 2003 | Compression effects on color and texture based multimedia indexing and retrievalabstractThis paper presents an evaluation of digital compression effects on content-based multimedia retrieval using color and texture attributes. Subjective evaluation tests that are applied on digital image and video databases using different compression and visual feature extraction techniques have been performed and reported. Simulations show that a satisfactory retrieval performance can be obtained from the compressed databases with 10% compression quality (i.e. 97.6% compression ratio in JPEG). Image retrieval based on HSV color histogram performs better than retrieval based on YUV color histogram in the uncompressed domain, and the other way around in the compressed domain. In general, video retrieval based on color histogram in MPEG-4 compressed databases performs better compared to H.263+ compressed databases. However, retrieval performance from H.263+ compressed databases at lower bit rates is more stable, where it drastically decreases in MPEG-4 compressed databases below 128 Kb/s. Retrieval based on texture features produces more robust performance than retrieval based on color. Subjective tests show that 25% compression quality achieves high compression ratio without loosing significant retrieval performance. The results are particularly relevant to applications in which a mobile device is involved in a multimedia retrieval system. Esin Guldogan, Olcay Guldogan, Serkan Kiranyaz, Kerem Caglar, Moncef Gabbouj |
ICIP (2) | 5 |
| 2003 | Random access using isolated regionsabstractRandom access is a desirable feature in many video communication systems. Intra pictures is conventionally used as random access points, but correct picture content is recovered gradually within a range of pictures starting from a non-intra random access point. This paper proposes the use of the isolated regions technique for gradual decoder refresh and presents how the proposed method can be used in the upcoming ITU-T recommendation H.264, also known as MPEG-4 part 10 or advanced video coding. The presented simulations reveal that the proposed method outperforms intra-picture-based random access points in error-prone network conditions. It is also shown that the proposed method is more flexible and suits packet-based transmission better compared to progressively located intra-coded slices. Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj |
ICIP (3) | 3 |
| 2002 | Coding of faded scene transitionsabstractCoding of a scene transition is often a challenging problem, from the compression efficiency point of view, because motion compensation may not be a powerful enough method to represent changes between pictures in the transition. This paper proposes a overlay coding technique for coding faded scene transitions. As shown by extensive simulations, over 50% bit-rate savings in both cross-fades and through-black fades compared to earlier techniques can be achieved. Overlay coding suits situations where video is edited manually or automatically. Dong Tian, Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj |
ICIP (2) | 4 |
| 2002 | Sub-picture: ROI coding and unequal error protectionabstractRegion-of-interest coding and unequal error protection are two important tools in video communication systems to improve the received visual quality. One common property of the two techniques is that unequal coding or transmission is applied to improve the quality of the most important parts of images. The proposed sub-picture coding technique facilitates both region-of-interest coding and unequal error protection by partitioning images to regions of interest and separating the corresponding coded data units from each other. Simulation results show that the overall subjective quality is considerably improved compared to the conventional coding schemes. Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj |
ICIP (3) | 3 |
| 2002 | The error concealment feature in the H.26L test modelabstractThis paper presents the error concealment (EC) feature implemented by the authors in the test model of the draft ITU-T video coding standard H.26L. The selected EC algorithms are based on weighted pixel value averaging for INTRA. pictures and boundary-matching-based motion vector recovery for INTER pictures. The specific concealment strategy and some special methods, including handling of B-pictures, multiple reference frames and entire frame losses, are described. Both subjective and objective results are given based on simulations under Internet conditions. The feature was adopted and is now included in the latest H.26L reference software TML-9.0. Ye-Kui Wang, Miska M. Hannuksela, Viktor Varsa, Ari Hourunranta, Moncef Gabbouj |
ICIP (2) | 5 |
| 2002 | Adaptive fuzzy order statistics-rational hybrid filters for color image processing
Lazhar Khriji, Moncef Gabbouj |
Fuzzy Sets Syst. | 2 |
| 2002 | Wavelet-based corner detection technique using optimal scale
Azhar Quddus, Moncef Gabbouj |
Pattern Recognit. Lett. | 2 |
| 2001 | Complexity of the consistency problem for certain Post classesabstractThe complexity of the consistency problem for several important classes of Boolean functions is analyzed. The classes of functions under investigation are those which are closed under function composition or superposition. Several of these so-called Post classes are considered within the context of machine learning with an application to breast cancer diagnosis. The considered Post classes furnish a user-selectable measure of reliability. It is shown that for realistic situations which may arise in practice, the consistency problem for these classes of functions is polynomial-time solvable. Ilya Shmulevich, Moncef Gabbouj, Jaakko Astola |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2000 | Wavelet based corner detection using singular value decompositionabstractIn this paper we present a novel technique for wavelet-based corner detection using singular value decomposition (SVD). Here SVD facilitates the selection of global natural scale in the discrete wavelet transform. We define natural scale as the level associated with most prominent (dominant) eigenvalue. The eigenvector corresponding to the dominant eigenvalue is considered as the natural scale. The corners are detected at the locations corresponding to modulus maxima. Results show the suitability of the approach. Comparison with a recently proposed technique is also provided. Azhar Quddus, Moncef Gabbouj |
ICASSP | 2 |
| 2000 | A New Image Similarity Measure Based on Ordinal CorrelationabstractWe propose an ordinal-based measure of image similarity. This measure is based on a previously developed general framework for image correspondence and incorporates region-based spatial information. The measure is capable of taking into account differences between images at various scales. Several examples are presented and the measure is evaluated on a set of test images. Bogdan Cramariuc, Ilya Shmulevich, Moncef Gabbouj, Asko Makela |
ICIP | 3 |
| 2000 | A Weighted Distance Approach to Relevance FeedbackabstractContent-based image retrieval systems use low-level features like color and texture for image representation. Given these representations as feature vectors, similarity between images is measured by computing distances in the feature space. Unfortunately, these low-level features cannot always capture the high-level concept of similarity in human perception. Relevance feedback tries to improve the performance by allowing iterative retrievals where the feedback information from the user is incorporated into the database search. We present a weighted distance approach where the weights are the ratios of standard deviations of the feature values both for the whole database and also among the images selected as relevant by the user. The feedback is used for both independent and incremental updating of the weights and these weights are used to iteratively refine the effects of different features in the database search. Retrieval performance is evaluated using average precision and progress that are computed on a database of approximately 10,000 images and an average performance improvement of 19% is obtained after the first iteration. Selim Aksoy, Robert M. Haralick, Faouzi Alaya Cheikh, Moncef Gabbouj |
ICPR | 4 |
| 2000 | Directional-rational approach for color image enhancementabstractIn this paper, we present an unsharp masking-based approach for noise smoothing and edge enhancing in multichannel images. The proposed structure is similar to the conventional unsharp masking structure, however, the enhancement is allowed only in the direction of maximal change and the enhancement parameter is computed as a nonlinear function of the rate of change. The proposed scheme enhances the true details, limits the overshoot near sharp edges and attenuates noise in flat areas. Moreover the use of the control function eliminates the need for the subjective coefficient /spl lambda/ used in the conventional unsharp masking technique. Simulations results show that the processed image presents sharp edges which makes it more pleasant to the human eye. Moreover, the amount of noise in the image is clearly reduced. Faouzi Alaya Cheikh, Moncef Gabbouj |
ISCAS | 2 |
| 1999 | Motion field estimation by vector rational interpolation for error concealment purposesabstractA study on the use of vector rational interpolation for the estimation of erroneously received motion fields of an MPEG-2 coded video bitstream has been performed. Four different motion vector interpolation schemes have been examined using motion information from available top and bottom adjacent blocks since left or right neighbours are usually lost. The presented interpolation schemes are capable of adapting their behaviour according to neighbouring motion information. Simulation results prove the satisfactory performance of the novel nonlinear interpolation schemes and the success of their application to the concealment of predictively coded frames. The motion vector rational interpolation concealment method proves to be a fast method, thus adequate for real-time applications. Sofia Tsekeridou, Faouzi Alaya Cheikh, Moncef Gabbouj, Ioannis Pitas |
ICASSP | 3 |
| 1999 | A Dedicated Hardware System for a Class of Nonlinear Order Statistics Rational Hybrid Filters with Applications to Image ProcessingabstractA dedicated hardware system is developed for a recent class of nonlinear hybrid filters called Order Statistics-Rational Hybrid Filters (OSRHF). The performance of these filters is also studied and compared against three effective nonlinear filters from the literature. The application at hand is noise filtering in grey level images. The proposed hardware system uses a residue number system (RNS) to compute the numerator and the denominator of the rational filter. The resulting structure is suitable for direct implementation on FPGAs. Lazhar Khriji, Giuseppe Bernacchia, Moncef Gabbouj, Giovanni L. Sicuranza |
ICIP (2) | 3 |
| 1999 | Programmable Hardware Implementation for the Median-Rational Hybrid FiltersabstractThee Median-Rational Hybrid Filter (MRHF) has been recently introduced as a new class of nonlinear filters and successfully applied to image filtering problems. The main characteristics of the MRHF are the good noise attenuation property in smooth areas and the preservation of edges end details. In fact, in changing areas the noise attenuation is traded for a good response to the change. Moreover, the median filters effectively remove the impulsive noise while the rational filter performs well in relatively high SNR Gaussian contaminated environments. Another relevant characteristic is that MRHF's usually act on small windows and thus require a reduced number of operations, resulting in simple and fast filtering structures. In this paper we present a programmable hardware implementation for the MRHF's. Exploiting the features of the dynamic logic families it is possible to achieve high speed and compactness, while keeping the power dissipation very low. Lazhar Khriji, Giuseppe Bernacchia, Moncef Gabbouj, Giovanni L. Sicuranza |
ICIP (4) | 3 |
| 1999 | Nonlinear Interpolators for Old Movie RestorationabstractA nonlinear interpolator using a rational function filter is applied to the restoration of image sequence frames of digitized old movies. Samples to be interpolated are due to stationary and random defects. The interpolator is preceded by a defect localization algorithm. The performance of the proposed interpolator has been assessed on several sequences and compared with a classical morphological operator. The hardware implementation of the proposed rational interpolator is also considered. Simulations show that the interpolated frames with the proposed operator are free from blockiness and jaggedness which are very difficult to avoid when using linear operators. Lazhar Khriji, Moncef Gabbouj, Stefano Marsi, Giovanni Ramponi, Etienne Decencière |
ICIP (3) | 2 |
| 1999 | Vector median-rational hybrid filters for multichannel image processingabstractIn this letter, a new class of nonlinear filters called vector median-rational hybrid filters (VMRHFs) for multispectral image processing is introduced and applied to the color image filtering problem. These filters are based on rational functions (RFs) offering a number of advantages. First, a rational function is a universal approximator and a good extrapolator. Second, it can be trained by a linear adaptive algorithm. Third, it produces the best approximation (w.r.t. a given cost function) for some specific functions. The output is the result of a vector rational operation over the output of three subfilters, such as two vector median (VM) subfilters and one center weighted vector median filter (CWVMF). These filters exhibit desirable properties, such as edge and details preservation and accurate chromaticity estimation. Lazhar Khriji, Moncef Gabbouj |
IEEE Signal Process. Lett. | 2 |
| 1999 | lambda-M-S filters for image restoration applicationsabstractA new filtering architecture is proposed, generalizing some previously introduced multilevel median filters. An efficient design procedure for the new filtering architecture is demonstrated for image restoration application. Simulation results show a good noise rejection performance, combined with a fine detail preservation capability. Doina Petrescu, Ioan Tabus, Moncef Gabbouj |
IEEE Trans. Image Process. | 3 |
| 1998 | Median-Rational Hybrid FiltersabstractWe introduce a new class of nonlinear filters called median-rational hybrid filters (MRHF) based on rational functions (RF). The filter output is the result of a rational operation taking into account three sub-functions, such as two FIR or median sub-filters and one center weighted median filter (CWMF). The proposed MRHF filters have the inherent property that on smooth areas they provide good noise attenuation whereas on changing areas the noise attenuation is traded for good response to the change. The performance of the proposed filter is compared against widely known nonlinear filters such as: morphological signal adaptive median filters, stack filters, rank-order morphological filters and simple rational filters. It is shown that a significant subjective improvement in the restored image quality as well as a consistent reduction in the objectively measured mean absolute error and mean square error is obtained. Lazhar Khriji, Moncef Gabbouj |
ICIP (2) | 2 |
| 1998 | Parallel Marker-Based Image Segmentation with Watershed Transformation
Alina N. Moga, Moncef Gabbouj |
J. Parallel Distributed Comput. | 2 |
| 1998 | Parallel watershed transformation algorithms for image segmentation
Alina N. Moga, Bogdan Cramariuc, Moncef Gabbouj |
Parallel Comput. | 3 |
| 1998 | Robust B-spline image modeling with application to image processingabstractIn this correspondence, we present a new approach to two-dimensional (2-D) robust spline image smoothing based on the M-estimator algorithm. Unlike in other M-estimator based image processing algorithms, the new algorithm takes into consideration the spatial relations between picture elements. The contribution of the sample to the model depends not only on the current residual of that sample, but also on the neighboring residuals. A smoothing parameter is estimated separately for each processing window and it adapts to the local structure of the image. The proposed algorithm is applied to image filtering. The resulting filter preserves details and suppresses additive Gaussian and impulsive noise efficiently. Marta Karczewicz, Moncef Gabbouj |
IEEE Trans. Image Process. | 2 |
| 1997 | Prediction based on Boolean, FIR-Boolean hybrid and stack filters for lossless image codingabstractThis paper proposes the use of mean absolute error (MAE) optimal Boolean and stack filters for sequential prediction in lossless grey-level image coding. FIR-Boolean hybrid filters are introduced as variations of Boolean filter structure and shown to be very effective for the prediction task. Different instances of optimal filtering are considered for realizing the prediction stage. First, the use of global-optimal predictors is analyzed, when the global MAE-optimal filter is used as a predictor. Then more refined structures, block-optimal and adaptive-size-block-optimal are considered, where predictors are adapted to local characteristics. These structures prove most suitable when small prediction masks are used. Extensive simulations are carried out for analyzing and comparing the performance of the newly introduced predictors and various other sequential predictors. Doina Petrescu, Ioan Tabus, Moncef Gabbouj |
ICASSP | 3 |
| 1997 | Prediction Based on Boolean Filters for Multiresolution Lossless Image CompressionabstractIn this paper Boolean filters and a variation of these, FIR-Boolean hybrid filters are proposed for realizing the prediction stages of a multiresolution lossless image compression structure. Optimal and adaptive Boolean filters are used for prediction and the proposed predictors performance is compared to the performance of other lossless multiresolution methods: the hierarchical interpolation scheme (HINT) and the S+P transform. Doina Petrescu, Moncef Gabbouj |
ICIP (2) | 2 |
| 1997 | Parallel Image Component Labeling With Watershed TransformationabstractThe parallel watershed transformation used in gray scale image segmentation is reconsidered on the basis of the component labeling problem. The main idea is to break the sequentiality of the watershed transformation and to correctly delimit the extent of all connected components locally, on each processor, simultaneously. The internal fragmentation of the catchment basins, due to domain decomposition, into smaller subcomponents is finally solved by employing a global connected components operator. Therefore, in a pyramidal structure of master-slave processors, internal contours of adjacent subcomponents within the same component are hierarchically removed. Global final connected areas are efficiently obtained in log/sub 2/ N steps on a logical grid of N processors. Timings and segmentation results of the algorithm built on top of the message passing interface and tested on the Gray T3D are brought forward to justify the superiority of the novel design solution compared against previous implementations. Alina N. Moga, Moncef Gabbouj |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | ECG data compression by spline approximation
Marta Karczewicz, Moncef Gabbouj |
Signal Process. | 2 |
| 1996 | Impulse noise removal in highly corrupted color imagesabstractWe present a novel and efficient technique for the restoration of color images which are highly corrupted with impulse noise. This is a detection-estimation based approach in which outliers are first detected using a Teager-like operator followed by a locally adaptive threshold. Center pixels whose "energy" exceeds some threshold are replaced with the local marginal median. Simulation results show the superior performance of the proposed filtering algorithm compared to the renowned vector median (VM) and generalized vector directional filter (GVDF), which are commonly used for color image restoration. Monte Carlo simulations show the edge preservation and impulse noise attenuation capabilities of the proposed technique. The efficiency of the algorithm stems from its simple arithmetic operations compared with more demanding ones, e.g. computation of distances and angles in the case of VMF and GVDF, respectively. Faouzi Alaya Cheikh, Ridha Hamila, Moncef Gabbouj, Jaakko Astola |
ICIP (1) | 3 |
| 1996 | A parallel marker based watershed transformationabstractThe parallel watershed transformation used in grayscale image segmentation is reconsidered on the basis of markers. The goal is to reduce the typical over segmentation by decreasing the number of catchment basins produced by flooding. Assimilating the set of basins with a weighted neighborhood graph and computing the minimum spanning forest in which every tree is rooted at a marked vertex, all non-marked regions in each tree are incorporated in the root region of the tree. A log/sub 2/N distributed message passing algorithm performing the above stated goal on N processors is presented. Two merits of the parallel algorithm are worth of mentioning: first, the local detection of the catchment basins conforming to the watershed principle (which strongly depends on the history of the region growth), with an extremely low communication traffic; and second, the parallel computation of the Boruvka like minimum spanning forest with the constraint that any tree contains exactly one marker. Evaluation of a Cray T3D implementation under the message passing interface (MPI) is included. Alina N. Moga, Moncef Gabbouj |
ICIP (2) | 2 |
| 1996 | Training based optimal stack filter design under structural constraintsabstractWe develop a new procedure for the optimal stack filter design under structural constraints, e.g. for minimizing an error criterion and simultaneously preserving the shape of some signals. The training framework for optimal stack filter design developed by Tabus, Petrescu and Gabbouj (see IEEE Transactions on Image Processing, Special Issue on Nonlinear Image Processing, IP-5, p.1-18, June 1996) provides us with a proper setting for imposing structural constraints to the optimal filter, in order to preserve some desired details of the image. The target application is optimal stack filter design for image restoration, the goal being the "most efficient" noise attenuation. Ioan Tabus, Doina Petrescu, Moncef Gabbouj |
ICIP (1) | 3 |
| 1996 | Order statistics learning vector quantizerabstractWe propose a novel class of learning vector quantizers (LVQs) based on multivariate data ordering principles. A special case of the novel LVQ class is the median LVQ, which uses either the marginal median or the vector median as a multivariate estimator of location. The performance of the proposed marginal median LVQ in color image quantization is demonstrated by experiments. Ioannis Pitas, Constantine Kotropoulos, Nikos Nikolaidis 0001, Ruikang Yang, Moncef Gabbouj |
IEEE Trans. Image Process. | 5 |
| 1996 | A training framework for stack and Boolean filtering-fast optimal design procedures and robustness case studyabstractA training framework is developed in this paper to design optimal nonlinear filters for various signal and image processing tasks. The targeted families of nonlinear filters are the Boolean filters and stack filters. The main merit of this framework at the implementation level is perhaps the absence of constraining models, making it nearly universal in terms of application areas. We develop fast procedures to design optimal or close to optimal filters, based on some representative training set. Furthermore, the training framework shows explicitly the essential part of the initial specification and how it affects the resulting optimal solution. Symmetry constraints are imposed on the data and, consequently, on the resulting optimal solutions for improved performance and ease of implementation. The case study is dedicated to natural images. The properties of optimal Boolean and stack filters, when the desired signal in the training set is the image of a natural scene, are analyzed. Specifically, the effect of changing the desired signal (using various natural images) and the characteristics of the noise (the probability distribution function, the mean, and the variance) is analyzed. Elaborate experimental conditions were selected to investigate the robustness of the optimal solutions using a sensitivity measure computed on data sets. A remarkably low sensitivity and, consequently, a good generalization power of Boolean and stack filters are revealed. Boolean-based filters are thus shown to be not only suitable for image restoration but also robust, making it possible to build libraries of "optimal" filters, which are suitable for a set of applications. Ioan Tabus, Doina Petrescu, Moncef Gabbouj |
IEEE Trans. Image Process. | 3 |
| 1995 | Fast order-recursive algorithms for optimal stack filter designabstractFor the nonlinear classes of filters based on the median archetype, e.g. stack, weighted order statistics (WOS), morphological filters the techniques exploiting recursiveness of embedded structures were not used yet. We investigate in this paper the possibilities of improving the speed of the optimal design techniques restating the optimal design problem as a sequence of optimal problems for embedded structures. The speed up effect of recursive-in-order design techniques is very significant but it is not the only benefit of this principle: it also allows the evaluation of the optimal structure of the filter, comparing the performances of the optimal filters having various structural parameters and observing where the performance index curve starts to flatten with the increase in the structure "size". Ioan Tabus, Moncef Gabbouj |
ICASSP | 2 |
| 1995 | An efficient watershed segmentation algorithm suitable for parallel implementationabstractAn important aspect of designing a parallel algorithm is exploitation of the data locality for minimization of the communication overhead. We propose a reformulation of a global image operation called the watershed transformation. The method is one of the various approaches for image segmentation and works by labeling connected components. Both serial and parallel programming models are presented and evaluated when running on SUN and DEC Alpha AXP workstations, and a Cray T3D, respectively. Alina N. Moga, Bogdan Cramariuc, Moncef Gabbouj |
ICIP | 3 |
| 1995 | Fast algorithms for analyzing and designing weighted median filters
Ruikang Yang, Moncef Gabbouj, Yrjö Neuvo |
Signal Process. | 2 |
| 1994 | Robust B-Spline Image SmoothingabstractIn this work we present a new approach to two-dimensional robust spline smoothing. The proposed method is based on M-estimator algorithms but unlike in other M-estimator based image processing algorithms it takes into consideration spatial relations between picture elements. The contribution of the sample to the model depends not only on the current residual of that sample, but also on the neighboring residuals. The smoothing parameter (/spl lambda/) is estimated separately for each processing window and it adapts to the local structure of the image. In order to test the proposed algorithm we apply it to image filtering problem. We show that the filter based on our algorithm has excellent detail preserving properties while suppressing additive Gaussian and impulsive noise very efficiently.> Marta Karczewicz, Moncef Gabbouj, Jaakko Astola |
ICIP (2) | 2 |
| 1994 | A Class of Order Statistics Learning Vector QuantizersabstractA novel class of Learning Vector Quantizers (LVQs) based on multivariate order statistics is proposed in order to overcome the drawback that the estimators for obtaining the reference vectors in LVQ do not have robustness either against erroneous choices for the winner vector or against the outliers that may exist in vector-valued observations. The performance of the proposed variants of LVQ is demonstrated by experiments. In the case of marginal median LVQ, its asymptotic properties are derived as well.> Ioannis Pitas, Constantine Kotropoulos, Nikos Nikolaidis 0001, Ruikang Yang, Moncef Gabbouj |
ISCAS | 5 |
| 1994 | The PBF of One Weight Weighted Median FiltersabstractThe general form of positive Boolean functions (PBFs) corresponding to one weight weighted median (WM) filters, which are a subclass of WM filters, is derived in this paper. Based on the results, it is straightforward to derive the corresponding PBF from a given one weight WM filter or vice versa. The results can also be used to check whether an PBF corresponds to an one weight WM filters. The checking procedure is very simple and tractable due to the fact that linear programming is not needed.> Tong Sun 0002, Moncef Gabbouj, Yrjö Neuvo |
ISCAS | 2 |
| 1994 | An Efficient Design Method for Optimal Weighted Median FilteringabstractEarlier research has shown that the problem of optimal weighted median filtering with structural constraints can be formulated as a nonconvex nonlinear programming problem in general. However, its high computational complexity and poor performance due to its nonconvex nature prohibit it from practical applications. In this paper, we shall show that the design problem can be formulated as a convex quadratic programming problem. The new algorithm is very efficient in the sense of computational complexity. The algorithm is also efficient in the sense of its capability to approach the global minimum. Using the algorithm optimal 1-D weighted median filters preserving pulses of length 3, 4 and 5 are tabulated.> Ruikang Yang, Moncef Gabbouj, Yrjö Neuvo |
ISCAS | 2 |
| 1994 | Parametric Analysis of Weighted Order Statistic FiltersabstractIn this paper, we study the convergence properties of weighted order statistic filters. Weighted order statistics filters are divided into five categories making their convergence properties easily understand. It is shown that any symmetric weighted order statistics filters will make the input sequence converge to a root or oscillate in a cycle of period 2. This result is significant since a restriction posed by earlier research is eliminated making the result applicable for the whole class of symmetric weighted order statistics filters. A condition to guarantee convergence of symmetric weighted order statistics filters is derived.> Ruikang Yang, Moncef Gabbouj, Pao-Ta Yu |
ISCAS | 2 |
| 1994 | Preface
Moncef Gabbouj |
Signal Process. | 1 |
| 1994 | Adaptive L-filters with applications in signal and image processing
Tong Sun 0002, Moncef Gabbouj, Yrjö Neuvo |
Signal Process. | 2 |
| 1994 | Center weighted median filters: Some properties and their applications in image processing
Tong Sun 0002, Moncef Gabbouj, Yrjö Neuvo |
Signal Process. | 2 |
| 1994 | Parametric analysis of weighted order statistics filtersabstractThe authors study the convergence properties of weighted order statistics filters. Based on a set of parameters, weighted order statistics filters are divided into five categories making their convergence properties easily understood. They show that any symmetric weighted order statistics filters will make the input sequence converge to a root or oscillate in a cycle of period 2. This result is significant since a restriction imposed by an earlier research is eliminated making the result applicable for the whole class of symmetric weighted order statistics filters. A condition to guarantee convergence of symmetric weighted order statistics filters is also derived.> Ruikang Yang, Moncef Gabbouj, Pao-Ta Yu |
IEEE Signal Process. Lett. | 2 |
| 1993 | Optimal weighted median filters under structural constraints
Ruikang Yang, Lin Yin, Moncef Gabbouj, Jaakko Astola, Yrjö Neuvo |
ISCAS | 3 |
| 1993 | Weighted medians - positive Boolean functions conversion algorithms
Jacek Nieweglowski, Moncef Gabbouj, Yrjö Neuvo |
Signal Process. | 2 |
| 1993 | Root properties of morphological filters
Qiaofei Wang, Moncef Gabbouj, Yrjö Neuvo |
Signal Process. | 2 |
| 1991 | Optimal stack filtering and classical Bayes decisionabstractOptimal stack filtering under the mean absolute error (MAE) criterion is studied. It is first shown that this problem is equivalent to the classical a priori Bayes minimum-cost decision. Generally, a linear program (LP) with O(b2/sup b/) variables and constraints (b is the window width) is required for finding the best filter. Instead, the authors develop a suboptimal routine which renders the use of the LP obsolete, but yields reasonably good filters. Sufficient conditions under which the proposed routine results in optimal solutions are provided and shown to hold in most practical cases. Several design examples are given.> Bing Zeng 0001, Moncef Gabbouj, Yrjö Neuvo |
ICASSP | 2 |
| 1991 | Speech production and speech modelling: Proceedings of the NATO Advanced Study Institute on Speech Production and Speech Modelling, Bonas, France, 17-29 July 1989, edited by William J. Hardcastle, Department of Linguistic Science, University of Reading, Reading, UK and Alain Marchal, CNRS, Aix-en-Provence, France, NATO ASI Series D: Behavioral and Social Sciences - Volume 55. Publishers: Kluwer Academic Publishers, P.O. Box 17, 3300 AA Dordrecht, The Netherlands, 1990, xi+448 pp., ISBN 0-7923-0746-1
Moncef Gabbouj |
Signal Process. | 1 |
| 1990 | Minimax optimization over the class of stack filtersabstractA new optimization theory for stack filters is presented in this paper. This new theory is based on the minimax error criterion rather than the mean absolute error (MAE) criterion used in [8]. In the binary case, a methodology will be designed to find the stack filter that minimizes the maximum absolute error between the input and the output signals. The most interesting feature of this optimization procedure is the fact that it can be solved using a linear program (LP), just like in the MAE case [8]. One drawback of this procedure is the problem of randomization due to the lost of structure in the constraint matrix of the LP. Several sub-optimal solutions will be discussed and an algorithm to find an optimal integer solution (still using a LP) under certain conditions will be provided. When generalizing to multiple-level inputs, complexity problems will arise and two alternatives will be suggested. One of these approaches assumes a parameterized stochastic model for the noise process and the LP is to pick the stack filter which minimizes the worst effect of the noise on the input signal. Moncef Gabbouj, Edward J. Coyle |
VCIP | 1 |
| 1990 | "Les filtres numériques, analyse et synthèse des filtres unidimensionnels", 3e édition, Collections techniques et scientifique des télécommunications (CNET-ENST): by R. Boite et H. Leich, Faculté Polytechnique de Mons. Publishers: Masson S.A., 120, bd Saint-Germain, 75280 Paris Cedex 06, France, 1990, 431 pages, ISBN 2-225-81884-3
Moncef Gabbouj |
Signal Process. | 1 |