Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Cagdas Bilen

dblp:99/4281 · also Çagdas Bilen · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 50% Multimedia analysis and retrieval · 50%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval
audio-visual analysis
0.912025
Learning to Highlight Audio by Watching Movies · CVPR 2025
Machine learning › Generative modeling
multimodal generation
0.312025
Learning to Highlight Audio by Watching Movies · CVPR 2025

Methods — techniques the papers use, named apart from their topics

transformer-based multimodal framework · 1.7pseudo-data generation · 0.9pseudo data generation · 0.9
YearPublicationVenuePosition
2026 Sound Event Detection With Boundary-Aware Optimization and Inference
abstract
Temporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers—Recurrent Event Detection (RED) and Event Proposal Network (EPN)—which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes.
Florian Schmid, Chi Ian Tang, Sanjeel Parekh, Vamsi K. Ithapu, Juan Azcarreta, Giacomo Ferroni, Yijun Qian, Arnoldas Jasonas, Cosmin Frateanu, Camilla Clark, Gerhard Widmer, Cagdas Bilen
IEEE Signal Process. Lett.12
2025 Learning to Highlight Audio by Watching Movies
abstract
Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal viewpoint selection or post-editing, has been central to media production, its natural counterpart, audio, has not undergone equivalent advancements. This often results in a disconnect between visual and acoustic saliency. To bridge this gap, we introduce a novel task: visually-guided acoustic highlighting, which aims to transform audio to deliver appropriate highlighting effects guided by the accompanying video, ultimately creating a more harmonious audio-visual experience. We propose a flexible, transformer-based multimodal framework to solve this task. To train our model, we also introduce a new dataset—THE MUDDY MIX DATASET, leveraging the meticulous audio and video crafting found in movies, which provides a form of free supervision. We develop a pseudo-data generation process to simulate poorly mixed audio, mimicking real-world scenarios through a three-step process—separation, adjustment, and remixing. Our approach consistently outperforms several baselines in both quantitative and subjective evaluation. We also systematically study the impact of different types of contextual guidance and difficulty levels of the dataset. Our project page is here: https://wikichao.github.io/VisAH/.
Chao Huang 0033, Ruohan Gao, J. M. F. Tsang, Jan Kurcius, Cagdas Bilen, Chenliang Xu, Sanjeel Parekh
CVPR5
2025 Efficient Neural and Numerical Methods for High-QualityOnline Speech Spectrogram Inversion via Gradient Theorem
Andres Fernandez, Juan Azcarreta, Cagdas Bilen, Jesus Monge-Alvarez
INTERSPEECH3
2021 Improving Sound Event Detection Metrics: Insights from DCASE 2020
abstract
The ranking of sound event detection (SED) systems may be biased by assumptions inherent to evaluation criteria and to the choice of an operating point. This paper compares conventional event-based and segment-based criteria against the Polyphonic Sound Detection Score (PSDS)'s intersection-based criterion, over a selection of systems from DCASE 2020 Challenge Task 4. It shows that, by relying on collars, the conventional event-based criterion introduces different strictness levels depending on the length of the sound events, and that the segment-based criterion may lack precision and be application dependent. Alternatively, PSDS's intersection-based criterion overcomes the dependency of the evaluation on sound event duration and provides robustness to labelling subjectivity, by allowing valid detections of interrupted events. Furthermore, PSDS enhances the comparison of SED systems by measuring sound event modelling performance independently from the systems' operating points.
Giacomo Ferroni, Nicolas Turpault, Juan Azcarreta, Francesco Tuveri, Romain Serizel, Cagdas Bilen, Sacha Krstulovic
ICASSP6
2020 A Framework for the Robust Evaluation of Sound Event Detection
abstract
This work defines a new framework for performance evaluation of polyphonic sound event detection (SED) systems, which overcomes the limitations of the conventional collar-based event decisions, event F-scores and event error rates. The proposed framework introduces a definition of event detection that is more robust against labelling subjectivity. It also resorts to polyphonic receiver operating characteristic (ROC) curves to deliver more global insight into system performance than F1-scores, and proposes a reduction of these curves into a single polyphonic sound detection score (PSDS), which allows system comparison independently from operating points (OPs). The presented method also delivers better insight into data biases and classification stability across sound classes. Furthermore, it can be tuned to varying applications in order to match a variety of user experience requirements. The benefits of the proposed approach are demonstrated by re-evaluating the baseline and two of the top-performing systems from DCASE 2019 Task 4.
Cagdas Bilen, Giacomo Ferroni, Francesco Tuveri, Juan Azcarreta, Sacha Krstulovic
ICASSP1
2016 Automatic allocation of NTF components for user-guided audio source separation
abstract
Nonnegative matrix or tensor factorization is a very popular approach for audio source separation. One important problem in nonnegative tensor factorization (NTF) in the context of user-guided audio source separation is the necessity to manually assign the NTF components to audio sources in order to be able to enforce prior information on the sources during the estimation process. In this paper, two new approaches to NTF based source separation are proposed, which do not require any manual component assignment to the sources, but estimate the underlying assignment automatically. Both algorithms use the prior information on the source samples in the estimation process along with either a limit on the minimum number of components each source uses or with a restriction that each component is used by sparse number of sources. The proposed methods are shown to outperform the classic approach with a manual distribution of the components equally among the sources.
Cagdas Bilen, Alexey Ozerov, Patrick Pérez
ICASSP1
2016 Multichannel audio declipping
abstract
Audio declipping consists in recovering so-called clipped audio samples that are set to a maximum / minimum threshold. Many different approaches were proposed to solve this problem in case of singlechannel (mono) recordings. However, while most of audio recordings are multichannel nowadays, there is no method designed specifically for multichannel audio declipping, where the inter-channel correlations may be efficiently exploited for a better declipping result. In this work we propose for the first time such a multichannel audio declipping method. Our method is based on representing a multichannel audio recording as a convolutive mixture of several audio sources, and on modeling the source power spectrograms and mixing filters by nonnegative tensor factorization model and full-rank covariance matrices, respectively. A generalized expectation-maximization algorithm is proposed to estimate model parameters. It is shown experimentally that the proposed multichannel audio de-clipping algorithm outperforms in average and in most cases a state-of-the-art single-channel declipping algorithm applied to each channel independently.
Alexey Ozerov, Cagdas Bilen, Patrick Pérez
ICASSP2
2016 Supervised learning of low-rank transforms for image retrieval
abstract
In this paper we propose a new method to automatically select the rank of linear transforms during supervised learning. Our approach relies on a sparsity-enforcing element-wise soft-thresholding operation applied after the linear transform. This novel approach to supervised rank learning has the important advantage that it is very simple to implement and incurs no extra complexity relative to linear transform learning. Furthermore, we propose a simple Stochastic Gradient Descent (SGD) implementation suitable for large scale learning, where SGD solvers have established themselves as the default workhorse. We compare our method to various other metric learning techniques in the application of image retrieval. This is one of the remaining few areas where supervised learning of low-rank linear transforms has not been fully exploited. The main reason for this is the lack of adequate datasets that are large enough, and hence we further introduce a new dataset consisting of groups of matching images derived from Cable News Network (CNN) videos using geometric verification and manual selection to find matching frames with adequate variability.
Cagdas Bilen, Joaquin Zepeda, Patrick Pérez
ICIP1
2010 On compressed sensing in parallel MRI of cardiac perfusion using temporal wavelet and TV regularization
abstract
Imaging of cardiac perfusion with MR is a challenging area of research especially due to the motion of the heart and limited time of data acquisition. Compressed sensing is a popular signal estimation method recently adopted by researchers in MRI which can improve the spatial and/or temporal resolution of the acquired images by reducing the number of necessary samples for image reconstruction. This paper focuses on performance of temporal regularization with total variation and wavelets in compressed sensing. The impact of the choice of regularization parameters on the image quality and the temporal variation of intensity in region of interests (ROIs) are discussed. It is found that selecting the regularization parameter so as to optimize the quality of the reconstructed image sequence as a whole, leads to erroneous reconstruction of certain regions due to over regularization.
Cagdas Bilen, Ivan W. Selesnick, Yao Wang 0001, Ricardo Otazo, Leon Axel, Daniel K. Sodickson
ICASSP1
2007 End-to-end stereoscopic video streaming with content-adaptive rate and format control
Anil Aksay, Selen Pehlivan, Engin Kurutepe, Cagdas Bilen, Tanir Ozcelebi, Gozde Bozdagi Akar, M. Reha Civanlar, A. Murat Tekalp
Signal Process. Image Commun.4
2006 A Multi-View Video Codec Based on H.264
abstract
H.264 is the current state-of-the-art monoscopic video codec providing almost twice the coding efficiency with the same quality comparing the previous codecs. With the increasing interest in 3D TV, multi-view video sequences that are provided by multiple cameras capturing the three dimensional objects and/or scene are more widely used. Compressing multi-view sequences independently with H.264 (simulcast) is not efficient since the redundancy between the closer cameras is not exploited. In order to reduce these redundancies, we propose a multi-view video codec based on H.264 using disparity estimation/compensation as well as motion estimation/compensation. In order to effectively search for disparity/motion without increasing computational complexity, we modified the buffering structure of H.264 and implemented several referencing modes. Our results show that for closely located cameras, our codec outperforms simulcast H.264 coding. For sparsely located cameras, our method can still improve coding gain depending on the video characteristics.
Cagdas Bilen, Anil Aksay, Gozde Bozdagi Akar
ICIP1
2006 End-to-End Stereoscopic Video Streaming System
abstract
Today, stereoscopic and multi-view video are among the popular research areas in the multimedia world. In this study, we have designed and built a platform consisting of stereo-view capturing, real-time transmission and display. At the display stage, end users view video in 3D by using polarized glasses. Multi-view video is compressed in an efficient way by using multi-view video coding techniques and streamed using standard real-time transport protocols. The entire system is built by modifying available open source systems whenever possible. Receiver can view the content of the video built from multiple channels as mono or stereo depending on its display and bandwidth capabilities
Selen Pehlivan, Anil Aksay, Cagdas Bilen, Gozde Bozdagi Akar, M. Reha Civanlar
ICME3