EDBT 2026 Demo / reviewers in the wild / expert
Serkan Kiranyaz
dblp:89/6384 · also Mustafa Serkan Kiranyaz
· DBLP profile ↗
92ranked-venue papers
24as first author
29since 2021 · last 2025
0000-0003-1551-3397ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 14 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 9 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 3Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BRSR-OpGAN: Blind radar signal restoration using operational generative adversarial networkabstractMany studies on radar signal restoration in the literature focus on isolated restoration problems, such as denoising over a certain type of noise, while ignoring other types of artifacts. Additionally, these approaches usually assume a noisy environment with a limited set of fixed signal-to-noise ratio (SNR) levels. However, real-world radar signals are often corrupted by a blend of artifacts, including but not limited to unwanted echo, sensor noise, intentional jamming, and interference, each of which can vary in type, severity, and duration. This study introduces Blind Radar Signal Restoration using an Operational Generative Adversarial Network (BRSR-OpGAN), which uses a dual domain loss in the temporal and spectral domains. This approach is designed to improve the quality of radar signals, regardless of the diversity and intensity of the corruption. The BRSR-OpGAN utilizes 1D Operational GANs, which use a generative neuron model specifically optimized for blind restoration of corrupted radar signals. This approach leverages GANs' flexibility to adapt dynamically to a wide range of artifact characteristics. The proposed approach has been extensively evaluated using a well-established baseline and a newly curated extended dataset called the Blind Radar Signal Restoration (BRSR) dataset. This dataset was designed to simulate real-world conditions and includes a variety of artifacts, each varying in severity. The evaluation shows an average SNR improvement over 15.1 dB and 14.3 dB for the baseline and BRSR datasets, respectively. Finally, the proposed approach can be applied in real-time, even on resource-constrained platforms. This pilot study demonstrates the effectiveness of blind radar restoration in time-domain for real-world radar signals, achieving exceptional performance across various SNR values and artifact types. The BRSR-OpGAN method exhibits robust and computationally efficient restoration of real-world radar signals, significantly outperforming existing methods. Muhammad Uzair Zahid, Serkan Kiranyaz, Alper Yildirim, Moncef Gabbouj |
Neural Networks | 2 |
| 2024 | Refining Myocardial Infarction Detection: A Novel Multi-Modal Composite Kernel Strategy in One-Class ClassificationabstractEarly detection of myocardial infarction (MI), a critical condition arising from coronary artery disease (CAD), is vital to prevent further myocardial damage. This study introduces a novel method for early MI detection using a one-class classification (OCC) algorithm in echocardiography. Our study overcomes the challenge of limited echocardiography data availability by adopting a novel approach based on Multi-modal Subspace Support Vector Data Description. The proposed technique involves a specialized MI detection framework employing multi-view echocardiography incorporating a composite kernel in the non-linear projection trick, fusing Gaussian and Laplacian sigmoid functions. Additionally, we enhance the update strategy of the projection matrices by adapting maximization for both or one of the modalities in the optimization process. Our method boosts MI detection capability by efficiently transforming features extracted from echocardiography data into an optimized lower-dimensional subspace. The OCC model trained specifically on target class instances from the comprehensive HMC-QU dataset that includes multiple echocardiography views indicates a marked improvement in MI detection accuracy. Our findings reveal that our proposed multi-view approach achieves a geometric mean of 71.24%, signifying a substantial advancement in echocardiography-based MI diagnosis and offering more precise and efficient diagnostic tools. Muhammad Uzair Zahid, Aysen Degerli, Fahad Sohrab, Serkan Kiranyaz, Tahir Hamid, Rashid Mazhar, Moncef Gabbouj |
ICIP | 4 |
| 2024 | Restoration of magnetohydrodynamic-corrupted 12-lead electrocardiogram to enhance cardiac monitoring during magnetic resonance imaging
Sakib Mahmud, Muhammad E. H. Chowdhury, Moajjem Hossain Chowdhury, Abdulrahman Alqahtani, Zaid Bin Mahbub, Faycal Bensaali, Serkan Kiranyaz |
Eng. Appl. Artif. Intell. | 7 |
| 2024 | Restoration of motion-corrupted EEG signals using attention-guided operational CycleGAN
Sakib Mahmud, Muhammad E. H. Chowdhury, Serkan Kiranyaz, Nasser Al-Emadi, Anas M. Tahir, Md. Shafayet Hossain, Amith Khandakar, Somaya Al-Máadeed |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | A novel deep learning technique for morphology preserved fetal ECG extraction from mother ECG using 1D-CycleGAN
Promit Basak, A. H. M. Nazmus Sakib, Muhammad E. H. Chowdhury, Nasser Al-Emadi, Huseyin Cagatay Yalcin, Shona Pedersen, Sakib Mahmud, Serkan Kiranyaz, Somaya Al-Máadeed |
Expert Syst. Appl. | 8 |
| 2024 | Wearable wrist to finger photoplethysmogram translation through restoration using super operational neural networks based 1D-CycleGAN for enhancing cardiovascular monitoringabstractPhysiological signals, such as the Photoplethysmogram (PPG) collected through wearable devices, consistently encounter significant motion artifacts. Current signal processing techniques, and even state-of-the-art machine learning algorithms, frequently struggle to effectively restore the inherent bodily signals amidst the array of randomly generated distortions. This often leads to the modification or even the degradation of the underlying physiological information. To enhance heart rate estimation from wrist PPG (wPPG) signals, this study introduces the Translation Through Restoration GAN (TTR-GAN). TTR-GAN comprises cascaded dual-stage 1D Cycle Generative Adversarial Networks (1D-CycleGANs) constructed using Super-ONNs. In the first phase, corrupted wPPG waveforms are blindly restored using a 1D-CycleGAN-based restoration framework. Subsequently, in the second phase, the restored wPPG waveforms are translated into clean finger PPG (fPPG) signals through a 1D-CycleGAN-based signal-to-signal translation or synthesis framework. Both the restorer and translator GANs undergo independent evaluation using robust temporal, spectral, and clinical metrics. The application of the multipass restoration scheme to the wPPG signals resulted in significantly lower entropy compared to the raw wPPGs, indicating reduced irregularity. Using the proposed PRTX metric to evaluate the translational ability of the multichannel translator CycleGAN, we achieved a substantial improvement of 35.88% in wrist-to-finger PPG translation. The correlation between the pulse rate and pulse rate variations estimated from the generated fPPG signals and the heart rate and heart rate variability readings from the ground truth ECG improved by approximately 10.4% and 14.7%, respectively, when compared to the raw wPPG signals. The proposed TTR-GAN can be implemented in wearable devices to obtain reliable real-time cardiovascular data during daily activities. Sakib Mahmud, Muhammad E. H. Chowdhury, Serkan Kiranyaz, Malisha Islam Tapotee, Purnata Saha, Anas M. Tahir, Amith Khandakar, Abdulrahman Alqahtani |
Expert Syst. Appl. | 3 |
| 2024 | Operational Support Estimator NetworksabstractIn this work, we propose a novel approach called Operational Support Estimator Networks (OSENs) for the support estimation task. Support Estimation (SE) is defined as finding the locations of non-zero elements in sparse signals. By its very nature, the mapping between the measurement and sparse signal is a non-linear operation. Traditional support estimators rely on computationally expensive iterative signal recovery techniques to achieve such non-linearity. Contrary to the convolutional layers, the proposed OSEN approach consists of operational layers that can learn such complex non-linearities without the need for deep networks. In this way, the performance of non-iterative support estimation is greatly improved. Moreover, the operational layers comprise so-called generative super neurons with non-local kernels. The kernel location for each neuron/feature map is optimized jointly for the SE task during training. We evaluate the OSENs in three different applications: i. support estimation from Compressive Sensing (CS) measurements, ii. representation-based classification, and iii. learning-aided CS reconstruction where the output of OSENs is used as prior knowledge to the CS algorithm for enhanced reconstruction. Experimental results show that the proposed approach achieves computational efficiency and outperforms competing methods, especially at low measurement rates by significant margins. Mete Ahishali, Mehmet Yamac, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | R2C-GAN: Restore-to-Classify Generative Adversarial Networks for blind X-ray restoration and COVID-19 classificationabstractRestoration of poor-quality medical images with a blended set of artifacts plays a vital role in a reliable diagnosis. As a pioneer study in blind X-ray restoration, we propose a joint model for generic image restoration and classification: Restore-to-Classify Generative Adversarial Networks (R2C-GANs). This is the first generic restoration approach forming an Image-to-Image translation task from poor-quality having noisy, blurry, or over/under-exposed images to high-quality image domain where forward and inverse transformations are learned using unpaired training samples. Simultaneously, the joint classification preserves the diagnostic-related label during restoration. Each R2C-GAN is equipped with operational layers/neurons in a compact architecture. The proposed joint model successfully restores images while achieving state-of-the-art Coronavirus Disease 2019 (COVID-19) classification with above 90% in F1-Score. In qualitative analysis, the restoration performance is confirmed by medical doctors where 68% of the restored images are selected against the original images. We share the software implementation at https://github.com/meteahishali/R2C-GAN. Mete Ahishali, Aysen Degerli, Serkan Kiranyaz, Tahir Hamid, Rashid Mazhar, Moncef Gabbouj |
Pattern Recognit. | 3 |
| 2023 | D2DLive: Iterative live video streaming algorithm for D2D networks
Zina Chkirbene, Ridha Hamila, Aiman Erbad, Serkan Kiranyaz, Nasser Al-Emadi |
Comput. Networks | 4 |
| 2023 | RamanNet: a generalized neural network architecture for Raman spectrum analysisabstractAbstract Raman spectroscopy provides a vibrational profile of the molecules and thus can be used to uniquely identify different kinds of materials. This sort of molecule fingerprinting has thus led to the widespread application of Raman spectrum in various fields like medical diagnosis, forensics, mineralogy, bacteriology, virology, etc. Despite the recent rise in Raman spectra data volume, there has not been any significant effort in developing generalized machine learning methods targeted toward Raman spectra analysis. We examine, experiment, and evaluate existing methods and conjecture that neither current sequential models nor traditional machine learning models are satisfactorily sufficient to analyze Raman spectra. Both have their perks and pitfalls; therefore, we attempt to mix the best of both worlds and propose a novel network architecture RamanNet. RamanNet is immune to the invariance property in convolutional neural networks (CNNs) and at the same time better than traditional machine learning models for the inclusion of sparse connectivity. This has been achieved by incorporating shifted multi-layer perceptrons (MLP) at the earlier levels of the network to extract significant features across the entire spectrum, which are further refined by the inclusion of triplet loss in the hidden layers. Our experiments on 4 public datasets demonstrate superior performance over the much more complex state-of-the-art methods, and thus, RamanNet has the potential to become the de facto standard in Raman spectra data analysis. Nabil Ibtehaz, Muhammad E. H. Chowdhury, Amith Khandakar, Serkan Kiranyaz, Mohammad Sohel Rahman, Susu M. Zughaier |
Neural Comput. Appl. | 4 |
| 2023 | Representation based regression for object distance estimationabstractIn this study, we propose a novel approach to predict the distances of the detected objects in an observed scene. The proposed approach modifies the recently proposed Convolutional Support Estimator Networks (CSENs). CSENs are designed to compute a direct mapping for the Support Estimation (SE) task in a representation-based classification problem. We further propose and demonstrate that representation-based methods (sparse or collaborative representation) can be used in well-designed regression problems especially over scarce data. To the best of our knowledge, this is the first representation-based method proposed for performing a regression task by utilizing the modified CSENs; and hence, we name this novel approach as Representation-based Regression (RbR). The initial version of CSENs has a proxy mapping stage (i.e., a coarse estimation for the support set) that is required for the input. In this study, we improve the CSEN model by proposing Compressive Learning CSEN (CL-CSEN) that has the ability to jointly optimize the so-called proxy mapping stage along with convolutional layers. The experimental evaluations using the KITTI 3D Object Detection distance estimation dataset show that the proposed method can achieve a significantly improved distance estimation performance over all competing methods. Finally, the software implementations of the methods are publicly shared at https://github.com/meteahishali/CSENDistance. Mete Ahishali, Mehmet Yamac, Serkan Kiranyaz, Moncef Gabbouj |
Neural Networks | 3 |
| 2023 | Generalized Tensor Summation Compressive Sensing Network (GTSNET): An Easy to Learn Compressive Sensing OperationabstractThe efforts in compressive sensing (CS) literature can be divided into two groups: finding a measurement matrix that preserves the compressed information at its maximum level, and finding a robust reconstruction algorithm. In the traditional CS setup, the measurement matrices are selected as random matrices, and optimization-based iterative solutions are used to recover the signals. Using random matrices when handling large or multi-dimensional signals is cumbersome especially when it comes to iterative optimizations. Recent deep learning-based solutions increase reconstruction accuracy while speeding up recovery, but jointly learning the whole measurement matrix remains challenging. For this reason, state-of-the-art deep learning CS solutions such as convolutional compressive sensing network (CSNET) use block-wise CS schemes to facilitate learning. In this work, we introduce a separable multi-linear learning of the CS matrix by representing the measurement signal as the summation of the arbitrary number of tensors. As compared to block-wise CS, tensorial learning eases blocking artifacts and improves performance, especially at low measurement rates (MRs), such as [Formula: see text]. The software implementation of the proposed network is publicly shared at https://github.com/mehmetyamac/GTSNET. Mehmet Yamac, Ugur Akpinar, Erdem Sahin, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Image Process. | 4 |
| 2023 | Robust Peak Detection for Holter ECGs by Self-Organized Operational Neural NetworksabstractAlthough numerous R-peak detectors have been proposed in the literature, their robustness and performance levels may significantly deteriorate in low-quality and noisy signals acquired from mobile electrocardiogram (ECG) sensors, such as Holter monitors. Recently, this issue has been addressed by deep 1-D convolutional neural networks (CNNs) that have achieved state-of-the-art performance levels in Holter monitors; however, they pose a high complexity level that requires special parallelized hardware setup for real-time processing. On the other hand, their performance deteriorates when a compact network configuration is used instead. This is an expected outcome as recent studies have demonstrated that the learning performance of CNNs is limited due to their strictly homogenous configuration with the sole linear neuron model. This has been addressed by operational neural networks (ONNs) with their heterogenous network configuration encapsulating neurons with various nonlinear operators. In this study, to further boost the peak detection performance along with an elegant computational efficiency, we propose 1-D Self-Organized ONNs (Self-ONNs) with generative neurons. The most crucial advantage of 1-D Self-ONNs over the ONNs is their self-organization capability that voids the need to search for the best operator set per neuron since each generative neuron has the ability to create the optimal operator during training. The experimental results over the China Physiological Signal Challenge-2020 (CPSC) dataset with more than one million ECG beats show that the proposed 1-D Self-ONNs can significantly surpass the state-of-the-art deep CNN with less computational complexity. Results demonstrate that the proposed solution achieves a 99.10% F1-score, 99.79% sensitivity, and 98.42% positive predictivity in the CPSC dataset, which is the best R-peak detection performance ever achieved. Moncef Gabbouj, Serkan Kiranyaz, Junaid Malik, Muhammad Uzair Zahid, Turker Ince, Muhammad E. H. Chowdhury, Amith Khandakar, Anas M. Tahir |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Convolutional Sparse Support Estimator Network (CSEN): From Energy-Efficient Support Estimation to Learning-Aided Compressive SensingabstractSupport estimation (SE) of a sparse signal refers to finding the location indices of the nonzero elements in a sparse representation. Most of the traditional approaches dealing with SE problems are iterative algorithms based on greedy methods or optimization techniques. Indeed, a vast majority of them use sparse signal recovery (SR) techniques to obtain support sets instead of directly mapping the nonzero locations from denser measurements (e.g., compressively sensed measurements). This study proposes a novel approach for learning such a mapping from a training set. To accomplish this objective, the convolutional sparse support estimator networks (CSENs), each with a compact configuration, are designed. The proposed CSEN can be a crucial tool for the following scenarios: 1) real-time and low-cost SE can be applied in any mobile and low-power edge device for anomaly localization, simultaneous face recognition, and so on and 2) CSEN's output can directly be used as "prior information," which improves the performance of sparse SR algorithms. The results over the benchmark datasets show that state-of-the-art performance levels can be achieved by the proposed approach with a significantly reduced computational complexity. Mehmet Yamac, Mete Ahishali, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | SRL-SOA: Self-Representation Learning with Sparse 1D-Operational Autoencoder for Hyperspectral Image Band SelectionabstractThe band selection in the hyperspectral image (HSI) data processing is an important task considering its effect on the computational complexity and accuracy. In this work, we propose a novel framework for the band selection problem: Self-Representation Learning (SRL) with Sparse 1D-Operational Autoencoder (SOA). The proposed SLR-SOA approach introduces a novel autoencoder model, SOA, that is designed to learn a representation domain where the data are sparsely represented. Moreover, the network composes of 1D-operational layers with the non-linear neuron model. Hence, the learning capability of neurons (filters) is greatly improved with shallow architectures. Using compact architectures is especially crucial in autoencoders as they tend to overfit easily because of their identity mapping objective. Overall, we show that the proposed SRL-SOA band selection approach outperforms the competing methods over two HSI data including Indian Pines and Salinas-A considering the achieved land cover classification accuracies. The software implementation of the SRL-SOA approach is shared publicly1. Mete Ahishali, Serkan Kiranyaz, Iftikhar Ahmad 0001, Moncef Gabbouj |
ICIP | 2 |
| 2022 | Osegnet: Operational Segmentation Network for Covid-19 Detection Using Chest X-Ray ImagesabstractCoronavirus disease 2019 (COVID-19) has been diagnosed automatically using Machine Learning algorithms over chest X-ray (CXR) images. However, most of the earlier studies used Deep Learning models over scarce datasets bearing the risk of overfitting. Additionally, previous studies have revealed the fact that deep networks are not reliable for classification since their decisions may originate from irrelevant areas on the CXRs. Therefore, in this study, we propose Operational Segmentation Network (OSegNet) that performs detection by segmenting COVID-19 pneumonia for a reliable diagnosis. To address the data scarcity encountered in training and especially in evaluation, this study extends the largest COVID-19 CXR dataset: QaTa-COV19 with 121,378 CXRs including 9258 COVID-19 samples with their corresponding ground-truth segmentation masks that are publicly shared with the research community. Consequently, OSegNet has achieved a detection performance with the highest accuracy of 99.65% among the state-of-the-art deep models with 98.09% precision. Aysen Degerli, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
ICIP | 2 |
| 2022 | Improved Domain Adaptation Approach for Bearing Fault DiagnosisabstractApplication of domain adaptation techniques to predictive maintenance of modern electric rotating machinery (RM) has significant potential with the goal of transferring or adaptation of a fault diagnosis model developed for one machine to be generalized on new machines and/or new working conditions. The generalized nonlinear extension of conventional convolutional neural networks (CNNs), the self-organized operational neural networks (Self-ONNs) are known to enhance the learning capability of CNN by introducing non-linear neuron models and further heterogeneity in the network configuration. In this study, first the state-of-the-art 1D CNNs and Self-ONNs are tested for cross-domain performance. Then, we propose to utilize Self-ONNs as feature extractor in the well-known domain-adversarial neural networks (DANN) to enhance its domain adaptation performance. Experimental results over the benchmark Case Western Reserve University (CWRU) real vibration data set for bearing fault diagnosis across different load domains demonstrate the effectiveness and feasibility of the proposed domain adaptation approach with similar computational complexity. Turker Ince, Sertac Kilickaya, Levent Eren, Ozer Can Devecioglu, Serkan Kiranyaz, Moncef Gabbouj |
IECON | 5 |
| 2022 | Federated Learning in NOMA Networks: Convergence, Energy and Fairness-Based DesignabstractFederated Learning (FL) is a collaborative machine learning (ML) approach, where different nodes in a network contribute to learning the model parameters. In addition, FL provides several attractive features such as data privacy and energy efficiency. Due to its collaborative nature, model parameters among nodes should be efficiently exchanged, while considering the scarce availability of clean spectral slots. In this work, we propose low-power efficient algorithms for FL of model parameters updates. We consider mobile edge nodes connected to a leading node (LD) with practical wireless links, where uplink updates from the nodes to the LD are shared without orthogonalizing the resources. In particular, we adopt a non-orthogonal multiple access (NOMA) uplink scheme, and investigate its effect on the convergence round (CR) of the model updates. Through deriving an analytical expression of the CR, we leverage it to formulate an optimization problem to minimize the total number of communication rounds and maximize the communication fairness among the nodes. We further investigate the performance of our proposed algorithms by considering different factors, including limited per-node energy and node heterogeneity. Monte-Carlo simulations are used to verify the accuracy of our derived expression of the CR. Moreover, through comprehensive simulation, we show that our proposed schemes largely reduce the communication latency between the LD and the nodes, and improve the communication fairness among the nodes. Ilyes Mrad, Lutfi Samara, Abubakr O. Al-Abbasi, Ridha Hamila, Aiman Erbad, Serkan Kiranyaz |
PIMRC | 6 |
| 2022 | Fully automated 2D and 3D convolutional neural networks pipeline for video segmentation and myocardial infarction detection in echocardiography
Oumaima Hamila, Sheela Ramanna, Christopher J. Henry, Serkan Kiranyaz, Ridha Hamila, Rashid Mazhar, Tahir Hamid |
Multim. Tools Appl. | 4 |
| 2021 | Reliable Covid-19 Detection using Chest X-Ray ImagesabstractCoronavirus disease 2019 (COVID-19) has emerged the need for computer-aided diagnosis with automatic, accurate, and fast algorithms. Recent studies have applied Machine Learning algorithms for COVID-19 diagnosis over chest X-ray (CXR) images. However, the data scarcity in these studies prevents a reliable evaluation with the potential of overfitting and limits the performance of deep networks. Moreover, these networks can discriminate COVID-19 pneumonia usually from healthy subjects only or occasionally, from limited pneumonia types. Thus, there is a need for a robust and accurate COVID-19 detector evaluated over a large CXR dataset. To address this need, in this study, we propose a reliable COVID-19 detection network: ReCovNet, which can discriminate COVID-19 pneumonia from 14 different thoracic diseases and healthy subjects. To accomplish this, we have compiled the largest COVID-19 CXR dataset: QaTa-COV19 with 124,616 images including 4603 COVID-19 samples. The proposed ReCovNet achieved a detection performance with 98.57% sensitivity and 99.77% specificity. Aysen Degerli, Mete Ahishali, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
ICIP | 3 |
| 2021 | Self-Organized Residual Blocks For Image Super-ResolutionabstractIt has become a standard practice to use the convolutional networks (ConvNet) with RELU non-linearity in image restoration and super-resolution (SR). Although the universal approximation theorem states that a multi-layer neural network can approximate any non-linear function with the desired precision, it does not reveal the best network architecture to do so. Recently, operational neural networks (ONNs) that choose the best non-linearity from a set of alternatives, and their “self-organized” variants (Self-ONN) that approximate any non-linearity via Taylor series have been proposed to address the well-known limitations and drawbacks of conventional ConvNets such as network homogeneity using only the McCulloch-Pitts neuron model. In this paper, we propose the concept of self-organized operational residual (SOR) blocks, and present hybrid network architectures combining regular residual and SOR blocks to strike a balance between the benefits of stronger non-linearity and the overall number of parameters. The experimental results demonstrate that the proposed architectures yield performance improvements in both PSNR and perceptual metrics. Onur Keles, A. Murat Tekalp, Junaid Malik, Serkan Kiranyaz |
ICIP | 4 |
| 2021 | Bm3d Vs 2-Layer OnnabstractDespite their recent success on image denoising, the need for deep and complex architectures still hinders the practical usage of CNNs. Older but computationally more efficient methods such as BM3D remain a popular choice, especially in resource-constrained scenarios. In this study, we aim to find out whether compact neural networks can learn to produce competitive results as compared to BM3D for AWGN image denoising. To this end, we conFigure networks with only two hidden layers and employ different neuron models and layer widths for comparing the performance with BM3D across different AWGN noise levels. Our results conclusively show that the recently proposed self-organized variant of operational neural networks based on a generative neuron model (Self-ONNs) is not only a better choice as compared to CNNs, but also provide competitive results as compared to BM3D and even significantly surpass it for high noise levels. Junaid Malik, Serkan Kiranyaz, Mehmet Yamac, Moncef Gabbouj |
ICIP | 2 |
| 2021 | Self-Organized Variational Autoencoders (Self-Vae) For Learned Image CompressionabstractIn end-to-end optimized learned image compression, it is standard practice to use a convolutional variational autoencoder with generalized divisive normalization (GDN) to transform images into a latent space. Recently, Operational Neural Networks (ONNs) that learn the best non-linearity from a set of alternatives, and their “self-organized” variants, Self-ONNs, that approximate any non-linearity via Taylor series have been proposed to address the limitations of convolutional layers and a fixed nonlinear activation. In this paper, we propose to replace the convolutional and GDN layers in the variational autoencoder with self-organized operational layers, and propose a novel self-organized variational autoencoder (Self-VAE) architecture that benefits from stronger non-linearity. The experimental results demonstrate that the proposed Self-VAE yields improvements in both rate-distortion performance and perceptual image quality. Mustafa Akin Yilmaz, Onur Keles, Hilal Güven, A. Murat Tekalp, Junaid Malik, Serkan Kiranyaz |
ICIP | 6 |
| 2021 | Speech Command Recognition in Computationally Constrained Environments with a Quadratic Self-Organized Operational LayerabstractAutomatic classification of speech commands has revolutionized human computer interactions in robotic applications. However, employed recognition models usually follow the methodology of deep learning with complicated networks which are memory and energy hungry. So, there is a need to either squeeze these complicated models or use more efficient lightweight models in order to be able to implement the resulting classifiers on embedded devices. In this paper, we pick the second approach and propose a network layer to enhance the speech command recognition capability of a lightweight network and demonstrate the result via experiments. The employed method borrows the ideas of Taylor expansion and quadratic forms to construct a better representation of features in both input and hidden layers. This richer representation results in recognition accuracy improvement as shown by extensive experiments on Google speech commands (GSC) and synthetic speech commands (SSC) datasets. Mohammad Soltanian, Junaid Malik, Jenni Raitoharju, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
IJCNN | 5 |
| 2021 | Cooperative Machine Learning Techniques for Cloud Intrusion DetectionabstractCloud computing is attracting a lot of attention in the past few years. Although, even with its wide acceptance, cloud security is still one of the most essential concerns of cloud computing. Many systems have been proposed to protect the cloud from attacks using attack signatures. Most of them may seem effective and efficient; however, there are many drawbacks such as the attack detection performance and the system maintenance. Recently, learning-based methods for security applications have been proposed for cloud anomaly detection especially with the advents of machine learning techniques. However, most researchers do not consider the attack classification which is an important parameter for proposing an appropriate countermeasure for each attack type. In this paper, we propose a new firewall model called Secure Packet Classifier (SPC) for cloud anomalies detection and classification. The proposed model is constructed based on collaborative filtering using two machine learning algorithms to gain the advantages of both learning schemes. This strategy increases the learning performance and the system's accuracy. To generate our results, a publicly available dataset is used for training and testing the performance of the proposed SPC. Our results show that the accuracy of the SPC model increases the detection accuracy by 20% compared to the existing machine learning algorithms while keeping a high attack detection rate. Zina Chkirbene, Ridha Hamila, Aiman Erbad, Serkan Kiranyaz, Nasser Al-Emadi, Mounir Hamdi |
IWCMC | 4 |
| 2021 | Exploiting heterogeneity in operational neural networks by synaptic plasticityabstractAbstract The recently proposed network model, Operational Neural Networks (ONNs), can generalize the conventional Convolutional Neural Networks (CNNs) that are homogenous only with a linear neuron model. As a heterogenous network model, ONNs are based on a generalized neuron model that can encapsulateanyset of non-linear operators to boost diversity and to learn highly complex and multi-modal functions or spaces with minimal network complexity and training data. However, the default search method to find optimal operators in ONNs, the so-called Greedy Iterative Search (GIS) method, usually takes several training sessions to find a single operator set per layer. This is not only computationally demanding, also the network heterogeneity is limited since the same set of operators will then be used for all neurons in each layer. To address this deficiency and exploit a superior level of heterogeneity, in this study the focus is drawn on searching the best-possible operator set(s) for the hidden neurons of the network based on the “Synaptic Plasticity” paradigm that poses the essential learning theory in biological neurons. During training, each operator set in the library can be evaluated by their synaptic plasticity level, ranked from the worst to the best, and an “elite” ONN can then be configured using the top-ranked operator sets found at each hidden layer. Experimental results over highly challenging problems demonstrate that the elite ONNs even with few neurons and layers can achieve a superior learning performance than GIS-based ONNs and as a result, the performance gap over the CNNs further widens. Serkan Kiranyaz, Junaid Malik, Habib Ben Abdallah, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Comput. Appl. | 1 |
| 2021 | Self-organized Operational Neural Networks with Generative NeuronsabstractOperational Neural Networks (ONNs) have recently been proposed to address the well-known limitations and drawbacks of conventional Convolutional Neural Networks (CNNs) such as network homogeneity with the sole linear neuron model. ONNs are heterogeneous networks with a generalized neuron model. However the operator search method in ONNs is not only computationally demanding, but the network heterogeneity is also limited since the same set of operators will then be used for all neurons in each layer. Moreover, the performance of ONNs directly depends on the operator set library used, which introduces a certain risk of performance degradation especially when the optimal operator set required for a particular task is missing from the library. In order to address these issues and achieve an ultimate heterogeneity level to boost the network diversity along with computational efficiency, in this study we propose Self-organized ONNs (Self-ONNs) with generative neurons that can adapt (optimize) the nodal operator of each connection during the training process. Moreover, this ability voids the need of having a fixed operator set library and the prior operator search within the library in order to find the best possible set of operators. We further formulate the training method to back-propagate the error through the operational layers of Self-ONNs. Experimental results over four challenging problems demonstrate the superior learning capability and computational efficiency of Self-ONNs over conventional ONNs and CNNs. Serkan Kiranyaz, Junaid Malik, Habib Ben Abdallah, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Networks | 1 |
| 2021 | Self-organized operational neural networks for severe image restoration problemsabstractDiscriminative learning based on convolutional neural networks (CNNs) aims to perform image restoration by learning from training examples of noisy-clean image pairs. It has become the go-to methodology for tackling image restoration and has outperformed the traditional non-local class of methods. However, the top-performing networks are generally composed of many convolutional layers and hundreds of neurons, with trainable parameters in excess of several million. We claim that this is due to the inherently linear nature of convolution-based transformation, which is inadequate for handling severe restoration problems. Recently, a non-linear generalization of CNNs, called the operational neural networks (ONN), has been shown to outperform CNN on AWGN denoising. However, its formulation is burdened by a fixed collection of well-known non-linear operators and an exhaustive search to find the best possible configuration for a given architecture, whose efficacy is further limited by a fixed output layer operator assignment. In this study, we leverage the Taylor series-based function approximation to propose a self-organizing variant of ONNs, Self-ONNs, for image restoration, which synthesizes novel nodal transformations on-the-fly as part of the learning process, thus eliminating the need for redundant training runs for operator search. In addition, it enables a finer level of operator heterogeneity by diversifying individual connections of the receptive fields and weights. We perform a series of extensive ablation experiments across three severe image restoration tasks. Even when a strict equivalence of learnable parameters is imposed, Self-ONNs surpass CNNs by a considerable margin across all problems, improving the generalization performance by up to 3 dB in terms of PSNR. Junaid Malik, Serkan Kiranyaz, Moncef Gabbouj |
Neural Networks | 2 |
| 2021 | Convolutional Sparse Support Estimator-Based COVID-19 Recognition From X-Ray ImagesabstractCoronavirus disease (COVID-19) has been the main agenda of the whole world ever since it came into sight. X-ray imaging is a common and easily accessible tool that has great potential for COVID-19 diagnosis and prognosis. Deep learning techniques can generally provide state-of-the-art performance in many classification tasks when trained properly over large data sets. However, data scarcity can be a crucial obstacle when using them for COVID-19 detection. Alternative approaches such as representation-based classification [collaborative or sparse representation (SR)] might provide satisfactory performance with limited size data sets, but they generally fall short in performance or speed compared to the neural network (NN)-based methods. To address this deficiency, convolution support estimation network (CSEN) has recently been proposed as a bridge between representation-based and NN approaches by providing a noniterative real-time mapping from query sample to ideally SR coefficient support, which is critical information for class decision in representation-based techniques. The main premises of this study can be summarized as follows: 1) A benchmark X-ray data set, namely QaTa-Cov19, containing over 6200 X-ray images is created. The data set covering 462 X-ray images from COVID-19 patients along with three other classes; bacterial pneumonia, viral pneumonia, and normal. 2) The proposed CSEN-based classification scheme equipped with feature extraction from state-of-the-art deep NN solution for X-ray images, CheXNet, achieves over 98% sensitivity and over 95% specificity for COVID-19 recognition directly from raw X-ray images when the average performance of 5-fold cross validation over QaTa-Cov19 data set is calculated. 3) Having such an elegant COVID-19 assistive diagnosis performance, this study further provides evidence that COVID-19 induces a unique pattern in X-rays that can be discriminated with high accuracy. Mehmet Yamac, Mete Ahishali, Aysen Degerli, Serkan Kiranyaz, Muhammad E. H. Chowdhury, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Generalized Operational Classifiers for Material IdentificationabstractMaterial is one of the intrinsic features of objects, and consequently material recognition plays an important role in image understanding. The same material may have various shapes and appearance, while keeping the same physical characteristic. This brings great challenges for material recognition. Besides suitable features, a powerful classifier also can improve the overall recognition performance. Due to the limitations of classical linear neurons, used in all shallow and deep neural networks, such as CNN, we propose to apply the generalized operational neurons to construct a classifier adaptively. These generalized operational perceptrons (GOP) contain a set of linear and nonlinear neurons, and possess a structure that can be built progressively. This makes GOP classifier more compact and can easily discriminate complex classes. The experiments demonstrate that GOP networks trained on a small portion of the data (4%) can achieve comparable performances to state-of-the-arts models trained on much larger portions of the dataset. Xiaoyue Jiang, Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Xiaoyi Feng |
MMSP | 4 |
| 2020 | Real-time phonocardiogram anomaly detection by adaptive 1D Convolutional Neural NetworksabstractThe heart sound signals (Phonocardiogram – PCG) enable the earliest monitoring to detect a potential cardiovascular pathology and have recently become a crucial tool as a diagnostic test in outpatient monitoring to assess heart hemodynamic status. The need for an automated and accurate anomaly detection method for PCG has thus become imminent. To determine the state-of-the-art PCG classification algorithm, 48 international teams competed in the PhysioNet (CinC) Challenge in 2016 over the largest benchmark dataset with 3126 records with the classification outputs, normal (N), abnormal (A) and unsure – too noisy (U). In this study, our aim is to push this frontier further; however, we focus deliberately on the anomaly detection problem while assuming a reasonably high Signal-to-Noise Ratio (SNR) on the records. By using 1D Convolutional Neural Networks trained with a novel data purification approach, we aim to achieve the highest detection performance and real-time processing ability with significantly lower delay and computational complexity. The experimental results over the high-quality subset of the same benchmark dataset show that the proposed approach achieves both objectives. Furthermore, our findings reveal the fact that further improvements indeed require a personalized (patient-specific) approach to avoid major drawbacks of a global PCG classification approach. Serkan Kiranyaz, Morteza Zabihi, Ali Bahrami Rad, Turker Ince, Ridha Hamila, Moncef Gabbouj |
Neurocomputing | 1 |
| 2020 | Progressive Operational Perceptrons with MemoryabstractGeneralized Operational Perceptron (GOP) was proposed to generalize the linear neuron model used in the traditional Multilayer Perceptron (MLP) by mimicking the synaptic connections of biological neurons showing nonlinear neurochemical behaviours. Previously, Progressive Operational Perceptron (POP) was proposed to train a multilayer network of GOPs which is formed layer-wise in a progressive manner. While achieving superior learning performance over other types of networks, POP has a high computational complexity. In this work, we propose POPfast, an improved variant of POP that signicantly reduces the computational complexity of POP, thus accelerating the training time of GOP networks. In addition, we also propose major architectural modications of POPfast that can augment the progressive learning process of POP by incorporating an information preserving, linear projection path from the input to the output layer at each progressive step. The proposed extensions can be interpreted as a mechanism that provides direct information extracted from the previously learned layers to the network, hence the term “memory”. This allows the network to learn deeper architectures and better data representations. An extensive set of experiments in human action, object, facial identity and scene recognition problems demonstrates that the proposed algorithms can train GOP networks much faster than POPs while achieving better performance compared to original POPs and other related algorithms. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
Neurocomputing | 2 |
| 2020 | Real-time throughput prediction for cognitive Wi-Fi networks
Muhammad Asif Khan 0001, Ridha Hamila, Nasser Al-Emadi, Serkan Kiranyaz, Moncef Gabbouj |
J. Netw. Comput. Appl. | 4 |
| 2020 | Operational neural networksabstractAbstract Feed-forward, fully connected artificial neural networks or the so-called multi-layer perceptrons are well-known universal approximators. However, their learning performance varies significantly depending on the function or the solution space that they attempt to approximate. This is mainly because of their homogenous configuration based solely on the linear neuron model. Therefore, while they learn very well those problems with a monotonous, relatively simple and linearly separable solution space, they may entirely fail to do so when the solution space is highly nonlinear and complex. Sharing the same linear neuron model with two additional constraints (local connections and weight sharing), this is also true for the conventional convolutional neural networks (CNNs) and it is, therefore, not surprising that in many challenging problems only the deep CNNs with a massive complexity and depth can achieve the required diversity and the learning performance. In order to address this drawback and also to accomplish a more generalized model over the convolutional neurons, this study proposes a novel network model, called operational neural networks (ONNs), which can be heterogeneous and encapsulate neurons with any set of operators to boost diversity and to learn highly complex and multi-modal functions or spaces with minimal network complexity and training data. Finally, the training method to back-propagate the error through the operational layers of ONNs is formulated. Experimental results over highly challenging problems demonstrate the superior learning capabilities of ONNs even with few neurons and hidden layers. Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neural Comput. Appl. | 1 |
| 2020 | Human experts vs. machines in taxa recognition
Johanna Ärje, Jenni Raitoharju, Alexandros Iosifidis, Ville Tirronen, Kristian Meissner, Moncef Gabbouj, Serkan Kiranyaz, Salme Kärkkäinen |
Signal Process. Image Commun. | 7 |
| 2020 | Patient-Specific Seizure Detection Using Nonlinear Dynamics and NullclinesabstractNonlinear dynamics has recently been extensively used to study epilepsy due to the complex nature of the neuronal systems. This study presents a novel method that characterizes the dynamic behavior of pediatric seizure events and introduces a systematic approach to locate the nullclines on the phase space when the governing differential equations are unknown. Nullclines represent the locus of points in the solution space where the components of the velocity vectors are zero. A simulation study over 5 benchmark nonlinear systems with well-known differential equations in three-dimensional exhibits the characterization efficiency and accuracy of the proposed approach that is solely based on the reconstructed solution trajectory. Due to their unique characteristics in the nonlinear dynamics of epilepsy, discriminative features can be extracted based on the nullclines concept. Using a limited training data (only 25% of each EEG record) in order to mimic the real-world clinical practice, the proposed approach achieves 91.15% average sensitivity and 95.16% average specificity over the benchmark CHB-MIT dataset. Together with an elegant computational efficiency, the proposed approach can, therefore, be an automatic and reliable solution for patient-specific seizure detection in long EEG recordings. Morteza Zabihi, Serkan Kiranyaz, Ville Jäntti, Tarmo Lipping, Moncef Gabbouj |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Heterogeneous Multilayer Generalized Operational PerceptronabstractThe traditional multilayer perceptron (MLP) using a McCulloch-Pitts neuron model is inherently limited to a set of neuronal activities, i.e., linear weighted sum followed by nonlinear thresholding step. Previously, generalized operational perceptron (GOP) was proposed to extend the conventional perceptron model by defining a diverse set of neuronal activities to imitate a generalized model of biological neurons. Together with GOP, a progressive operational perceptron (POP) algorithm was proposed to optimize a predefined template of multiple homogeneous layers in a layerwise manner. In this paper, we propose an efficient algorithm to learn a compact, fully heterogeneous multilayer network that allows each individual neuron, regardless of the layer, to have distinct characteristics. Based on the complexity of the problem, the proposed algorithm operates in a progressive manner on a neuronal level, searching for a compact topology, not only in terms of depth but also width, i.e., the number of neurons in each layer. The proposed algorithm is shown to outperform other related learning methods in extensive experiments on several classification problems. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | 1-D Convolutional Neural Networks for Signal Processing Applicationsabstract1D Convolutional Neural Networks (CNNs) have recently become the state-of-the-art technique for crucial signal processing applications such as patient-specific ECG classification, structural health monitoring, anomaly detection in power electronics circuitry and motor-fault detection. This is an expected outcome as there are numerous advantages of using an adaptive and compact 1D CNN instead of a conventional (2D) deep counterparts. First of all, compact 1D CNNs can be efficiently trained with a limited dataset of 1D signals while the 2D deep CNNs, besides requiring 1D to 2D data transformation, usually need datasets with massive size, e.g., in the "Big Data" scale in order to prevent the well-known "overfitting" problem. 1D CNNs can directly be applied to the raw signal (e.g., current, voltage, vibration, etc.) without requiring any pre- or post-processing such as feature extraction, selection, dimension reduction, denoising, etc. Furthermore, due to the simple and compact configuration of such adaptive 1D CNNs that perform only linear 1D convolutions (scalar multiplications and additions), a real-time and low-cost hardware implementation is feasible. This paper reviews the major signal processing applications of compact 1D CNNs with a brief theoretical background. We will present their state-of-the-art performances and conclude with focusing on some major properties. Serkan Kiranyaz, Turker Ince, Osama Abdeljaber, Onur Avci, Moncef Gabbouj |
ICASSP | 1 |
| 2019 | Knowledge Transfer for Face Verification Using Heterogeneous Generalized Operational PerceptronsabstractFace verification is a prominent biometric technique for identity authentication that has been used extensively in several security applications. In practice, face verification is often performed along with other visual surveillance tasks in the computing device. Thus, the ability to share the computation and reuse the information already extracted for other analysis tasks can greatly help reduce the computation load on the devices. In this study, we propose to utilize the knowledge transfer approach for the face verification problem by building a heterogeneous neural network architecture of Generalized Operational Perceptrons on top of the intermediate features extracted for object recognition purpose. Experimental results show that using our proposed approach, a face verification system can be incorporated into an existing visual analysis system with less additional memory and computational cost, compared to other similar approaches. Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
ICIP | 2 |
| 2019 | PyGOP: A Python library for Generalized Operational Perceptron algorithms
Dat Thanh Tran, Serkan Kiranyaz, Moncef Gabbouj, Alexandros Iosifidis |
Knowl. Based Syst. | 2 |
| 2018 | Acceleration Approaches for Big Data AnalysisabstractThe massive size of data that needs to be processed by Machine Learning models nowadays sets new challenges related to their computational complexity and memory footprint. These challenges span all processing steps involved in the application of the related models, i.e., from the fundamental processing steps needed to evaluate distances of vectors, to the optimization of large-scale systems, e.g. for non-linear regression using kernels, or the speed up of deep learning models formed by billions of parameters. In order to address these challenges, new approximate solutions have been recently proposed based on matrix/tensor decompositions, randomization and quantization strategies. This paper provides a comprehensive review of the related methodologies and discusses their connections. Anton Muravev, Dat Thanh Tran, Moncef Gabbouj, Alexandros Iosifidis, Serkan Kiranyaz |
ICIP | 5 |
| 2018 | 1-D CNNs for structural damage detection: Verification on a structural health monitoring benchmark data
Osama Abdeljaber, Onur Avci, Serkan Kiranyaz, Boualem Boashash, Henry Sodano, Daniel J. Inman |
Neurocomputing | 3 |
| 2018 | Benchmark database for fine-grained image classification of benthic macroinvertebrates
Jenni Raitoharju, Ekaterina Riabchenko, Iftikhar Ahmad 0001, Alexandros Iosifidis, Moncef Gabbouj, Serkan Kiranyaz, Ville Tirronen, Johanna Ärje, Salme Kärkkäinen, Kristian Meissner |
Image Vis. Comput. | 6 |
| 2018 | Feature synthesis for image classification and retrieval via one-against-all perceptrons
Jenni Raitoharju, Serkan Kiranyaz, Moncef Gabbouj |
Neural Comput. Appl. | 2 |
| 2018 | Spatiotemporal Saliency Estimation by Spectral Foreground DetectionabstractWe present a novel approach for spatiotemporal saliency detection by optimizing a unified criterion of color contrast, motion contrast, appearance, and background cues. To this end, we first abstract the video by temporal superpixels. Second, we propose a novel graph structure exploiting the saliency cues to assign the edge weights. The salient segments are then extracted by applying a spectral foreground detection method, quantum cuts, on this graph. We evaluate our approach on several public datasets for video saliency and activity localization to demonstrate the favorable performance of the proposed video quantum cuts compared to the state of the art. Çaglar Aytekin, Horst Possegger, Thomas Mauthner, Serkan Kiranyaz, Horst Bischof, Moncef Gabbouj |
IEEE Trans. Multim. | 4 |
| 2017 | A k-nearest neighbor multilabel ranking algorithm with application to content-based image retrievalabstractMultilabel ranking is an important machine learning task with many applications, such as content-based image retrieval (CBIR). However, when the number of labels is large, traditional algorithms are either infeasible or show poor performance. In this paper, we propose a simple yet effective multilabel ranking algorithm that is based on k-nearest neighbor paradigm. The proposed algorithm ranks labels according to the probabilities of the label association using the neighboring samples around a query sample. Different from traditional approaches, we take only positive samples into consideration and determine the model parameters by directly optimizing ranking loss measures. We evaluated the proposed algorithm using four popular multilabel datasets. The proposed algorithm achieves equivalent or better performance than other instance-based learning algorithms. When applied to a CBIR system with a dataset of 1 million samples and over 190 thousand labels, which is much larger than any other multilabel datasets used earlier, the proposed algorithm clearly outperforms the competing algorithms. Honglei Zhang 0001, Serkan Kiranyaz, Moncef Gabbouj |
ICASSP | 2 |
| 2017 | Generalized model of biological neural networks: Progressive operational perceptronsabstractTraditional Artificial Neural Networks (ANNs) such as Multi-Layer Perceptrons (MLPs) and Radial Basis Functions (RBFs) were designed to simulate biological neural networks; however, they are based only loosely on biology and only provide a crude model. This in turn yields well-known limitations and drawbacks on the performance and robustness. In this paper we shall address them by introducing a novel feed-forward ANN model, Generalized Operational Perceptrons (GOPs) that consist of neurons with distinct (non-)linear operators to achieve a generalized model of the biological neurons and ultimately a superior diversity. We modified the conventional back-propagation (BP) to train GOPs and furthermore, proposed Progressive Operational Perceptrons (POPs) to achieve self-organized and depth-adaptive GOPs according to the learning problem. The most crucial property of the POPs is their ability to simultaneously search for the optimal operator set and train each layer individually. The final POP is, therefore, formed layer by layer and this ability enables POPs with minimal network depth to attack the most challenging learning problems that cannot be learned by conventional ANNs even with a deeper and significantly complex configuration. Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
IJCNN | 1 |
| 2017 | The effect of automated taxa identification errors on biological indices
Johanna Ärje, Salme Kärkkäinen, Kristian Meissner, Alexandros Iosifidis, Turker Ince, Moncef Gabbouj, Serkan Kiranyaz |
Expert Syst. Appl. | 7 |
| 2017 | Progressive Operational Perceptrons
Serkan Kiranyaz, Turker Ince, Alexandros Iosifidis, Moncef Gabbouj |
Neurocomputing | 1 |
| 2017 | Extended quantum cuts for unsupervised salient object extraction
Çaglar Aytekin, Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
Multim. Tools Appl. | 3 |
| 2017 | Learning graph affinities for spectral graph-based salient object detection
Çaglar Aytekin, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
Pattern Recognit. | 3 |
| 2016 | Face segmentation in thumbnail images by data-adaptive convolutional segmentation networksabstractIn this study we address the problem of face segmentation in thumbnail images. While there have been several approaches for face detection, none performs detection in such low resolution and segmentation with pixel accuracy. In this paper, we propose convolutional segmentation networks (CSNs) that can be trained to learn segmentation of human faces. Unlike the deep classifiers such as Convolutional Neural Network (CNNs), CSNs have the unique design solely for segmentation with minimal complexity. Furthermore, we propose a self-data organization (SDO) in order to create “expert” CSNs each of which is specialized over a set of images with certain face characteristics. SDO is integrated with CSN training in an interleaved manner and it is the key for the learning with simple and compact networks rather than the deep ones. This is especially a desired property for the limited face datasets with challenging face variations and complexities. Evaluations on the benchmark dataset show that CSNs can achieve an elegant segmentation accuracy despite the limited training data size, thumbnail resolution and highly complex face modalities. Serkan Kiranyaz, Muhammad-Adeel Waris, Iftikhar Ahmad 0001, Ridha Hamila, Moncef Gabbouj |
ICIP | 1 |
| 2016 | Salient object segmentation based on linearly combined affinity graphsabstractIn this paper, we propose a graph affinity learning method for a recently proposed graph-based salient object detection method, namely Extended Quantum Cuts (EQCut). We exploit the fact that the output of EQCut is differentiable with respect to graph affinities, in order to optimize linear combination coefficients and parameters of several differentiable affinity functions by applying error backpropagation. We show that the learnt linear combination of affinities improves the performance over the baseline method and achieves comparable (or even better) performance when compared to the state-of-the-art salient object segmentation methods. Çaglar Aytekin, Alexandros Iosifidis, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 3 |
| 2016 | Joint K-Means quantization for Approximate Nearest Neighbor SearchabstractRecently, Approximate Nearest Neighbor (ANN) Search has become a very popular approach for similarity search on large-scale datasets. In this paper, we propose a novel vector quantization method for ANN, which introduces a joint multi-layer K-Means clustering solution for determination of the codebooks. The performance of the proposed method is improved further by a joint encoding scheme. Experimental results verify the success of the proposed algorithm as it outperforms the state-of-the-art methods. Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 2 |
| 2016 | Learned vs. engineered features for fine-grained classification of aquatic macroinvertebratesabstractAquatic macroinvertebrate biomonitoring is an efficient way of assessment of slow and subtle anthropogenic changes and their effect on water quality. It is imperative to have reliable identification and counts of the various taxa occurring in samples as these form the basis for the quality indices used to infer the ecological status of the aquatic ecosystem. In this paper, we try to close the gap between human taxa identification accuracy (typically 90-95% on 30-40 classes of macroinvertebrates) and results of automatic fine-grained classification by introducing a novel technique based on Convolutional Neural Networks (CNN). CNN learns optimal features for macroinvertebrate classification and achieves near human accuracy when tested on 29 macroinvertebrate classes. Moreover, we perform comparative evaluation of the learned features against the hand-crafted features, which have been commonly used in classical approaches, and confirm superiority of the learned deep features over the engineered ones. Ekaterina Riabchenko, Kristian Meissner, Iftikhar Ahmad 0001, Alexandros Iosifidis, Ville Tirronen, Moncef Gabbouj, Serkan Kiranyaz |
ICPR | 7 |
| 2016 | An Optimized k-NN Approach for Classification on Imbalanced Datasets with Missing Data
Ezgi C. Ozan, Ekaterina Riabchenko, Serkan Kiranyaz, Moncef Gabbouj |
IDA | 3 |
| 2016 | Learning to rank salient segments extracted by multispectral Quantum Cuts
Çaglar Aytekin, Serkan Kiranyaz, Moncef Gabbouj |
Pattern Recognit. Lett. | 2 |
| 2016 | K-Subspaces Quantization for Approximate Nearest Neighbor SearchabstractApproximate Nearest Neighbor (ANN) search has become a popular approach for performing fast and efficient retrieval on very large-scale datasets in recent years, as the size and dimension of data grow continuously. In this paper, we propose a novel vector quantization method for ANN search which enables faster and more accurate retrieval on publicly available datasets. We define vector quantization as a multiple affine subspace learning problem and explore the quantization centroids on multiple affine subspaces. We propose an iterative approach to minimize the quantization error in order to create a novel quantization scheme, which outperforms the state-of-the-art algorithms. The computational cost of our method is also comparable to that of the competing methods. Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Competitive Quantization for Approximate Nearest Neighbor SearchabstractIn this study, we propose a novel vector quantization algorithm for Approximate Nearest Neighbor (ANN) search, based on a joint competitive learning strategy and hence called as competitive quantization (CompQ). CompQ is a hierarchical algorithm, which iteratively minimizes the quantization error by jointly optimizing the codebooks in each layer, using a gradient decent approach. An extensive set of experimental results and comparative evaluations show that CompQ outperforms the-state-of-the-art while retaining a comparable computational complexity. Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Training Radial Basis Function Neural Networks for Classification via Class-Specific ClusteringabstractIn training radial basis function neural networks (RBFNNs), the locations of Gaussian neurons are commonly determined by clustering. Training inputs can be clustered on a fully unsupervised manner (input clustering), or some supervision can be introduced, for example, by concatenating the input vectors with weighted output vectors (input-output clustering). In this paper, we propose to apply clustering separately for each class (class-specific clustering). The idea has been used in some previous works, but without evaluating the benefits of the approach. We compare the class-specific, input, and input-output clustering approaches in terms of classification performance and computational efficiency when training RBFNNs. To accomplish this objective, we apply three different clustering algorithms and conduct experiments on 25 benchmark data sets. We show that the class-specific approach significantly reduces the overall complexity of the clustering, and our experimental results demonstrate that it can also lead to a significant gain in the classification performance, especially for the networks with a relatively few Gaussian neurons. Among other applied clustering algorithms, we combine, for the first time, a dynamic evolutionary optimization method, multidimensional particle swarm optimization, and the class-specific clustering to optimize the number of cluster centroids and their locations. Jenni Raitoharju, Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Visual saliency by extended quantum cutsabstractIn this study, we propose an unsupervised, state-of-the-art saliency map generation algorithm which is based on a recently proposed link between quantum mechanics and spectral graph clustering, Quantum Cuts. The proposed algorithm forms a graph among superpixels extracted from an image and optimizes a criterion related to the image boundary, local contrast and area information. Furthermore, the effects of the graph connectivity, superpixel shape irregularity, superpixel size and how to determine the affinity between superpixels are analyzed in detail. Furthermore, we introduce a novel approach to propose several saliency maps. Resulting saliency maps consistently achieves a state-of-the-art performance in a large number of publicly available benchmark datasets in this domain, containing around 18k images in total. Çaglar Aytekin, Ezgi C. Ozan, Serkan Kiranyaz, Moncef Gabbouj |
ICIP | 3 |
| 2015 | Long-term epileptic EEG classification via 2D mapping and textural features
Kaveh Samiee, Serkan Kiranyaz, Moncef Gabbouj, Tapio Saramäki |
Expert Syst. Appl. | 2 |
| 2014 | Automatic Object Segmentation by Quantum CutsabstractIn this study, the link between quantum mechanics and graph-cuts is exploited and a novel saliency map generation and salient object segmentation method is proposed based on the ground state solution of a modified Hamiltonian. First, the graph representation of certain quantum mechanical operators is studied. This reveals strong connections with widely used graph-cut algorithms while quantum mechanical constraints exhibit crucial advantages over the existing graph-cut algorithms. Furthermore, concepts such as potential field helps solving a particular singularity problem related to Laplacian matrices. In the proposed approach, the ground state (wave function) corresponding to a sub-atomic particle of a modified Hamiltonian operator corresponds to a particular optimization problem, the solution of which yields the salient object segmentation in a digital image. This approach provides a parameter-free -hence dataset independent-, unsupervised and fully automatic saliency map generation, which outperforms many existing state-of-the-art algorithms. The results of the proposed salient object extraction method exhibit such a promising accuracy that pushes the frontier in this field to the borders of the input-driven processing only - without the use of "object knowledge" aided by long-term human memory and intelligence. Furthermore, with the novel technologies for measuring a quantum wave function, the proposed method has a unique potential: Salient object segmentation in an actual physical setup in nano-scale. Such an unprece-dendent property will not only produce segmentation results instantaneously, but may be a unique opportunity to achieve accurate object segmentation in real-time for the massive visual repositories of today's "Big Data". Çaglar Aytekin, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 2 |
| 2014 | Incremental Learning with Support Vector Data DescriptionabstractDue to the simplicity and firm mathematical foundation, Support Vector Machines (SVMs) have been intensively used to solve classification problems. However, training SVMs on real world large-scale databases is computationally costly and sometimes infeasible when the dataset size is massive and non-stationary. In this paper, we propose an incremental learning approach that greatly reduces the time consumption and memory usage for training SVMs. The proposed method is fully dynamic, which stores only a small fraction of previous training examples whereas the rest can be discarded. It can further handle unseen labels in new training batches. The classification experiments show that the proposed method achieves the same level of classification accuracy as batch learning while the computational cost is significantly reduced, and it can outperform other incremental SVM approaches for the new class problem. Weiyi Xie, Stefan Uhlmann, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 3 |
| 2014 | Automated patient-specific classification of long-term Electroencephalography
Serkan Kiranyaz, Turker Ince, Morteza Zabihi, Dilek Ince |
J. Biomed. Informatics | 1 |
| 2014 | A perceptual scheme for fully automatic video shot boundary detection
Murat Birinci, Serkan Kiranyaz |
Signal Process. Image Commun. | 2 |
| 2014 | Integrating Color Features in Polarimetric SAR Image ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) data are used extensively for terrain classification applying SAR features from various target decompositions and certain textural features. However, one source of information has so far been neglected from PolSAR classification: Color. It is a common practice to visualize PolSAR data by color coding methods and thus, it is possible to extract powerful color features from such pseudocolor images so as to provide additional data for a superior terrain classification. In this paper, we first review previous attempts for PolSAR classifications using various feature combinations and then we introduce and perform in-depth investigation of the application of color features over the Pauli color-coded images besides SAR and texture features. We run an extensive set of comparative evaluations using 24 different feature set combinations over three images of the Flevoland- and the San Francisco Bay region from the RADARSAT-2 and the AIRSAR systems operating in C- and L-bands, respectively. We then consider support vector machines and random forests classifier topologies to test and evaluate the role of color features over the classification performance. The classification results show that the additional color features introduce a new level of discrimination and provide noteworthy improvement in classification performance (compared with the traditionally employed PolSAR and texture features) within the application of land use and land cover classification. Stefan Uhlmann, Serkan Kiranyaz |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2013 | Quantum mechanics in computer vision: Automatic object extractionabstractAn automatic object extraction method is proposed exploiting the rich mathematical structure of quantum mechanics. First, a novel segmentation method based on the solutions of Schrödinger's equation is proposed. This powerful segmentation method allows us to model complex objects and inherent structures of edge, shape, and texture information along with the grey-level intensity uniformity, all in a single equation. Due to the large amount of segments extracted with the proposed method, the selection of the object segment is performed by maximizing a regularization energy function based on a recently proposed sub-segment analysis indicating the object boundaries. The results of the proposed automatic object extraction method exhibit such a promising accuracy that pushes the frontier in this field to the borders of the input-driven processing only — without the use of “object knowledge” aided by long-term human memory and intelligence. Çaglar Aytekin, Serkan Kiranyaz, Moncef Gabbouj |
ICIP | 2 |
| 2013 | Evaluation of classifiers for polarimetric SAR classificationabstractPolarimetric SAR data is been extensively used for the application of land use and land cover classification. Various classifier approaches have been applied to many different polarimetric images employing numerous features. In this paper, we want to provide an evaluation of commonly used supervised classifiers within the field of polarimetric SAR classification considering the effects of different number of training samples. Two polarimetric SAR images are considered representing an easier 4 class and more complex 15 class problem using a small set of eigen-decomposition features and tested with Neural Network, SVM, and Decision Tree classifiers. Results show that already rather small training sets can provide comparable results reducing the need for large labeled training data especially considering more challenging classification tasks. This can be further investigated in the area of semi-supervised learning. Stefan Uhlmann, Serkan Kiranyaz |
IGARSS | 2 |
| 2013 | Polarimetric SAR classification using visual color features extracted over pseudo color imagesabstractPolarimetric SAR data have been used extensively for terrain classification applying primitives from various target decompositions as well as texture features. However, there is a source of information that has been neglected so far from polarimetric SAR classification: Color. It is a common practice to visualize polarimetric SAR data by color coding methods and thus it is possible to extract powerful color features from such pseudocolor images. In this paper, we investigate and evaluate discrimination power of color features extracted over various pseudocolor images. Experiments are conducted over the San Francisco Bay region on RADARSAT-2 data by using Support Vector Machines. The classification results show that the additional color features introduce a new level of discrimination and provide noteworthy improvement in classification performance (compared to the traditionally employed polarimetric SAR and texture features) within the application of land use and land cover classification. Stefan Uhlmann, Serkan Kiranyaz, Moncef Gabbouj |
IGARSS | 2 |
| 2012 | Evolutionary RBF classifier for polarimetric SAR images
Turker Ince, Serkan Kiranyaz, Moncef Gabbouj |
Expert Syst. Appl. | 2 |
| 2012 | Dynamic and scalable audio classification by collective network of binary classifiers framework: An evolutionary approach
Serkan Kiranyaz, Toni Mäkinen, Moncef Gabbouj |
Neural Networks | 1 |
| 2011 | Multi-dimensional evolutionary feature synthesis for content-based image retrievalabstractLow-level features (also called descriptors) play a central role in content-based image retrieval (CBIR) systems. Features are various types of information extracted from the content and represent some of its characteristics or signatures. However, especially the (low-level) features, which can be extracted automatically usually lack the discrimination power needed for accurate description of the image content and may lead to a poor retrieval performance. In order to efficiently address this problem, in this paper we propose a multi- dimensional evolutionary feature synthesis technique, which seeks for the optimal linear and non-linear operators so as to synthesize highly discriminative set of features in an optimal dimension. The optimality therein is sought by the multi-dimensional particle swarm optimization method along with the fractional global-best formation technique. Clustering and CBIR experiments where the proposed feature synthesizer is evolved using only the minority of the image database, demonstrate a significant performance improvement and exhibit a major discrimination between the features of different classes. Serkan Kiranyaz, Jenni Raitoharju, Turker Ince, Moncef Gabbouj |
ICIP | 1 |
| 2011 | Incremental evolution of collective network of binary classifier for polarimetric SAR image classificationabstractIn this paper, we propose a dedicated application of collective network of binary classifiers (CNBC) to address the problem of incremental learning, which occurs by introducing new SAR terrain classes. Furthermore, another major goal is to achieve a high classification performance over multiple SAR images even though the training data may not be entirely accurate. The CNBC in principle adopts a “Divide and Conquer” type approach by allocating an individual network of binary classifiers (NBCs) to discriminate each SAR terrain class among others and performing evolutionary search to find the optimal binary classifier (BC) in each NBC. Such design further allows dynamic SAR class and feature scalability in such a way that the CNBC can gradually adapt its internal topology to new features and classes with minimal effort. Experiments visually demonstrate the classification accuracy and efficiency of the proposed system over eight fully polarimetric NASA/JPL AIRSAR data sets. Stefan Uhlmann, Serkan Kiranyaz, Moncef Gabbouj, Turker Ince |
ICIP | 2 |
| 2011 | Multi-dimensional particle swarm optimization in dynamic environments
Serkan Kiranyaz, Jenni Raitoharju, Moncef Gabbouj |
Expert Syst. Appl. | 1 |
| 2011 | Personalized long-term ECG classification: A systematic approach
Serkan Kiranyaz, Turker Ince, Jenni Raitoharju, Moncef Gabbouj |
Expert Syst. Appl. | 1 |
| 2010 | Dynamic Data Clustering Using Stochastic Approximation Driven Multi-Dimensional Particle Swarm Optimization
Serkan Kiranyaz, Turker Ince, Moncef Gabbouj |
EvoApplications (1) | 1 |
| 2010 | Network of evolutionary binary classifiers for classification and retrieval in macroinvertebrate databasesabstractIn this paper, we focus on advanced classification and data retrieval schemes that are instrumental when processing large taxonomical image datasets. With large number of classes, classification and an efficient retrieval of a particular benthic macroinvertebrate image within a dataset will surely pose a severe problem. To address this, we propose a novel network of evolutionary binary classifiers, which is scalable, dynamically adaptable and highly accurate for the classification and retrieval of large biological species-image datasets. The classification and retrieval results for the macroinvertebrate test data attain taxonomic accuracy that equals and even surpasses that of an average expert. Our findings are encouraging for aquatic biomonitoring where cost intensity of sample analysis currently poses a bottleneck for routine biomonitoring. Serkan Kiranyaz, Moncef Gabbouj, Jenni Raitoharju, Turker Ince, Kristian Meissner |
ICIP | 1 |
| 2010 | Classification of Polarimetric SAR Images Using Evolutionary RBF NetworksabstractThis paper proposes an evolutionary RBF network classifier for polar metric synthetic aperture radar ( SAR) images. The proposed feature extraction process utilizes the full covariance matrix, the gray level co-occurrence matrix (GLCM) based texture features, and the backscattering power (Span) combined with the H/α/A decomposition, which are projected onto a lower dimensional feature space using principal component analysis. An experimental study is performed using the fully polar metric San Francisco Bay data set acquired by the NASA/Jet Propulsion Laboratory Airborne SAR (AIRSAR) at L-band to evaluate the performance of the proposed classifier. Classification results (in terms of confusion matrix, overall accuracy and classification map) compared to the Wish art and a recent NN-based classifiers demonstrate the effectiveness of the proposed algorithm. Turker Ince, Serkan Kiranyaz, Moncef Gabbouj |
ICPR | 2 |
| 2010 | Evaluation of global and local training techniques over feed-forward neural network architecture spaces for computer-aided medical diagnosis
Turker Ince, Serkan Kiranyaz, Jenni Raitoharju, Moncef Gabbouj |
Expert Syst. Appl. | 2 |
| 2010 | Perceptual color descriptor based on spatial distribution: A top-down approach
Serkan Kiranyaz, Murat Birinci, Moncef Gabbouj |
Image Vis. Comput. | 1 |
| 2010 | Fractional Particle Swarm Optimization in Multidimensional Search SpaceabstractIn this paper, we propose two novel techniques, which successfully address several major problems in the field of particle swarm optimization (PSO) and promise a significant breakthrough over complex multimodal optimization problems at high dimensions. The first one, which is the so-called multidimensional (MD) PSO, re-forms the native structure of swarm particles in such a way that they can make interdimensional passes with a dedicated dimensional PSO process. Therefore, in an MD search space, where the optimum dimension is unknown, swarm particles can seek both positional and dimensional optima. This eventually removes the necessity of setting a fixed dimension a priori, which is a common drawback for the family of swarm optimizers. Nevertheless, MD PSO is still susceptible to premature convergences due to lack of divergence. Among many PSO variants in the literature, none yields a robust solution, particularly over multimodal complex problems at high dimensions. To address this problem, we propose the fractional global best formation (FGBF) technique, which basically collects all the best dimensional components and fractionally creates an artificial global best (aGB) particle that has the potential to be a better "guide" than the PSO's native gbest particle. This way, the potential diversity that is present among the dimensions of swarm particles can be efficiently used within the aGB particle. We investigated both individual and mutual applications of the proposed techniques over the following two well-known domains: 1) nonlinear function minimization and 2) data clustering. An extensive set of experiments shows that in both application domains, MD PSO with FGBF exhibits an impressive speed gain and converges to the global optima at the true dimension regardless of the search space dimension, swarm size, and the complexity of the problem. Serkan Kiranyaz, Turker Ince, Alper Yildirim, Moncef Gabbouj |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2009 | Evolutionary artificial neural networks by multi-dimensional particle swarm optimization
Serkan Kiranyaz, Turker Ince, Alper Yildirim, Moncef Gabbouj |
Neural Networks | 1 |
| 2008 | Unsupervised design of Artificial Neural Networks via multi-dimensional Particle Swarm OptimizationabstractIn this paper, we present a novel and efficient approach for automatic design of artificial neural networks (ANNs) by evolving to the optimal network configuration(s) within an architecture space. The evolution technique, the so-called multidimensional particle swarm optimization (MD PSO) re-forms the native structure of PSO particles in such a way that they can make inter-dimensional passes with a dedicated dimensional PSO process. So in a multidimensional search space where the optimum dimension is unknown, swarm particles can seek for both positional and dimensional optima. This eventually removes the necessity of setting a fixed dimension a priori, which is a common drawback for the family of swarm optimizers. With the proper encoding of the network configurations and parameters into particles, MD PSO can then seek for positional optimum in the error space and dimensional optimum in the architecture space. The optimum dimension converged at the end of a MD PSO process corresponds to a unique ANN configuration where the network parameters (connections, weights and biases) can then be resolved from the positional optimum reached on that dimension. The efficiency and performance of the proposed technique is demonstrated over one of the hardest synthetic problems. The experimental results show that MD PSO evolves to optimum or near-optimum networks in general. Serkan Kiranyaz, Turker Ince, Alper Yildirim, Moncef Gabbouj |
ICPR | 1 |
| 2008 | A Generic Shape/Texture Descriptor Over Multiscale Edge Field: 2-D Walking Ant HistogramabstractA novel shape descriptor, which can be extracted from the major object edges automatically and used for the multimedia content-based retrieval in multimedia databases, is presented. By adopting a multiscale approach over the edge field where the scale represents the amount of simplification, the most relevant edge segments, referred to as subsegments, which eventually represent the major object boundaries, are extracted from a scale-map. Similar to the process of a walking ant with a limited line of sight over the boundary of a particular object, we traverse through each subsegment and describe a certain line of sight, whether it is a continuous branch or a corner, using individual 2-D histograms. Furthermore, the proposed method can also be tuned to be an efficient texture descriptor, which achieves a superior performance especially for directional textures. Finally, integrating the whole process as feature extraction module into MUVIS framework allows us to test the mutual performance of the proposed shape descriptor in the context of multimedia indexing and retrieval. Serkan Kiranyaz, Miguel Ferreira, Moncef Gabbouj |
IEEE Trans. Image Process. | 1 |
| 2007 | Hierarchical Cellular Tree: An Efficient Indexing Scheme for Content-Based Retrieval on Multimedia DatabasesabstractOne of the challenges in the development of a content-based multimedia indexing and retrieval application is to achieve an efficient indexing scheme. The developers and users who are accustomed to making queries to retrieve a particular multimedia item from a large scale database can be frustrated by the long query times. Conventional indexing structures cannot usually cope with the requirements of a multimedia database, such as dynamic indexing or the presence of high-dimensional audiovisual features. Such structures do not scale well with the ever increasing size of multimedia databases whilst inducing corruption and resulting in an over-crowded indexing structure. This paper addresses such problems and presents a novel indexing technique, hierarchical cellular tree (HCT), which is designed to bring an effective solution especially for indexing large multimedia databases. Furthermore it provides an enhanced browsing capability, which enables user to make a guided tour within the database. A pre-emptive cell-search mechanism is introduced in order to prevent corruption, which may occur due to erroneous item insertions. Among the hierarchical levels that are built in a bottom-up fashion, similar items are collected into appropriate cellular structures at some level. Cells are subject to mitosis operations when the dissimilarity exceeds a required level. By mitosis operations, cells are kept focused and compact and yet, they can grow into any dimension as long as the compactness is maintained. The proposed indexing scheme is then used along with a recently introduced query method, the progressive query, in order to achieve the ultimate goal, from the user point of view that is retrieval of the most relevant items in the earliest possible time regardless of the database size. Experimental results show that the speed of retrievals is significantly improved and the indexing structure shows no sign of degradations when the database size is increased. Furthermore, HCT indexing body can conveniently be used for efficient browsing and navigation operations among the multimedia database items Serkan Kiranyaz, Moncef Gabbouj |
IEEE Trans. Multim. | 1 |
| 2006 | Multi-Scale Edge Detection and Object Extraction for Image RetrievalabstractA new scheme for boundary-based object extraction and description of still images, through multi-scale edge detection, is proposed in this paper. Boundary-based methods try to extract closed contours from individual edge pixels through edge-linking. Our approach is based on a connected structure of edge pixels as the initial edge-linking elements. These connected structures, the sub-segments, are extracted from the Canny edge map of an image. Multiple simplification-scales are derived from applying iterations of the bilateral filter to the image, providing extra information about the relative importance of each sub-segment. Edge-linking towards contour closure is achieved through perceptually-driven minimum cost search. Furthermore, a shape-based description vector is derived from the extracted contours, and retrieval results are obtained via the integration of the whole scheme into MUVIS framework Miguel Ferreira, Serkan Kiranyaz, Moncef Gabbouj |
ICASSP (2) | 2 |
| 2006 | A generic audio classification and segmentation approach for multimedia indexing and retrievalabstractWe focus the attention on the area of generic and automatic audio classification and segmentation for audio-based multimedia indexing and retrieval applications. In particular, we present a fuzzy approach toward hierarchic audio classification and global segmentation framework based on automatic audio analysis providing robust, bi-modal, efficient and parameter invariant classification over global audio segments. The input audio is split into segments, which are classified as speech, music, fuzzy or silent. The proposed method minimizes critical errors of misclassification by fuzzy region modeling, thus increasing the efficiency of both pure and fuzzy classification. The experimental results show that the critical errors are minimized and the proposed framework significantly increases the efficiency and the accuracy of audio-based retrieval especially in large multimedia databases. Serkan Kiranyaz, Ahmad Farooq Qureshi, Moncef Gabbouj |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Automatic Object Extraction Over Multiscale Edge Field for Multimedia RetrievalabstractIn this work, we focus on automatic extraction of object boundaries from Canny edge field for the purpose of content-based indexing and retrieval over image and video databases. A multiscale approach is adopted where each successive scale provides further simplification of the image by removing more details, such as texture and noise, while keeping major edges. At each stage of the simplification, edges are extracted from the image and gathered in a scale-map, over which a perceptual subsegment analysis is performed in order to extract true object boundaries. The analysis is mainly motivated by Gestalt laws and our experimental results suggest a promising performance for main objects extraction, even for images with crowded textural edges and objects with color, texture, and illumination variations. Finally, integrating the whole process as feature extraction module into MUVIS framework allows us to test the mutual performance of the proposed object extraction method and subsequent shape description in the context of multimedia indexing and retrieval. A promising retrieval performance is achieved, and especially in some particular examples, the experimental results show that the proposed method presents such a retrieval performance that cannot be achieved by using other features such as color or texture. Serkan Kiranyaz, Miguel Ferreira, Moncef Gabbouj |
IEEE Trans. Image Process. | 1 |
| 2005 | A dynamic content-based indexing method for multimedia databases: hierarchical cellular treeabstractThis paper presents a novel indexing technique, hierarchical cellular tree, which is designed to bring an effective solution especially for indexing on large-scale multimedia databases. A pre-emptive cell search mechanism is introduced in order to prevent the corruption of large multimedia item collections due to the limited discrimination obtained from the visual and aural descriptors. In addition to this, the similar items are focused within appropriate cellular structures, which will be the subject to mitosis operations when the dissimilarity emerges as a result of irrelevant item insertions. Mitosis operations ensure to keep the cells in a focused and compact form and yet the cells can grow into any dimension as long as the compactness prevails. The proposed indexing scheme is then optimized for a novel query method, the progressive query, in order to maximize the retrieval efficiency for the user point of view. Experimental results show that the speed of the retrievals is significantly improved. Serkan Kiranyaz |
ICIP (1) | 1 |
| 2005 | An Efficient Image Retrieval Scheme on Java Enabled Mobile DevicesabstractContent-based image retrieval over wireless networks is a challenging research problem. In this paper, we present an efficient content-based image retrieval framework, which is developed for mobile platforms in client-server architecture and uses a combination of various low-level visual features. Several techniques were adapted in order to achieve the retrieval efficiency and query speed on mobile networks. Particularly a new implementation, which is called compact media retrieval on progressive query (CMR-PQ), is introduced. CMR-PQ is basically designed to retrieve the query results in an acceptable time, a network adaptive and configurable scheme. The experimental results present various query completion and image retrieval timings over several networks Iftikhar Ahmad 0001, Serkan Kiranyaz, Moncef Gabbouj |
MMSP | 2 |
| 2003 | Compression effects on color and texture based multimedia indexing and retrievalabstractThis paper presents an evaluation of digital compression effects on content-based multimedia retrieval using color and texture attributes. Subjective evaluation tests that are applied on digital image and video databases using different compression and visual feature extraction techniques have been performed and reported. Simulations show that a satisfactory retrieval performance can be obtained from the compressed databases with 10% compression quality (i.e. 97.6% compression ratio in JPEG). Image retrieval based on HSV color histogram performs better than retrieval based on YUV color histogram in the uncompressed domain, and the other way around in the compressed domain. In general, video retrieval based on color histogram in MPEG-4 compressed databases performs better compared to H.263+ compressed databases. However, retrieval performance from H.263+ compressed databases at lower bit rates is more stable, where it drastically decreases in MPEG-4 compressed databases below 128 Kb/s. Retrieval based on texture features produces more robust performance than retrieval based on color. Subjective tests show that 25% compression quality achieves high compression ratio without loosing significant retrieval performance. The results are particularly relevant to applications in which a mobile device is involved in a multimedia retrieval system. Esin Guldogan, Olcay Guldogan, Serkan Kiranyaz, Kerem Caglar, Moncef Gabbouj |
ICIP (2) | 3 |