EDBT 2026 Demo / reviewers in the wild / expert
A. Taylan Cemgil
dblp:41/6613 · also Ali Taylan Cemgil
· DBLP profile ↗
64ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0003-4463-8455ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 3 · 2 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
18 papers |
Trustworthy machine learning · 48% Representation and self-supervised learning · 10% Optimization for machine learning · 8% | |
| Computer graphics and multimedia
9 papers |
Audio and music processing · 98% Multimedia analysis and retrieval · 1% Image and video processing · 1% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 56% Medical and health informatics · 25% Computational science and engineering · 19% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 30 heaviest of 73, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
robustness |
1.7 | 4 | 2022 | Evaluating the Adversarial Robustness of Adaptive Test-time Defenses · ICML 2022 A Fine-Grained Analysis on Distribution Shift · ICLR 2022 Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Trustworthy machine learning › robustness
robust representations |
0.9 | 2 | 2020 | The Autoencoding Variational Autoencoder · NeurIPS 2020 Adversarially Robust Representations with Smooth Encoders · ICLR 2020 |
Machine learning › Trustworthy machine learning › fairness
bias evaluation |
0.8 | 1 | 2024 | Evaluating Model Bias Requires Characterizing its Mistakes · ICML 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Evaluating Model Bias Requires Characterizing its Mistakes · ICML 2024 |
Machine learning › Graph learning
graph neural network |
0.7 | 1 | 2023 | Transformers Meet Directed Graphs · ICML 2023 |
Machine learning › Graph learning › graph neural network
graph transformer |
0.7 | 1 | 2023 | Transformers Meet Directed Graphs · ICML 2023 |
Machine learning › Trustworthy machine learning › robustness › model robustness evaluation
adversarial robustness evaluation |
0.6 | 1 | 2022 | Evaluating the Adversarial Robustness of Adaptive Test-time Defenses · ICML 2022 |
Machine learning › Learning paradigms › supervised learning
classifier training |
0.6 | 1 | 2022 | Learning Optimal Conformal Classifiers · ICLR 2022 |
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction |
0.6 | 1 | 2022 | Learning Optimal Conformal Classifiers · ICLR 2022 |
Machine learning › Trustworthy machine learning › robustness
distribution shift |
0.6 | 1 | 2022 | A Fine-Grained Analysis on Distribution Shift · ICLR 2022 |
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification |
0.6 | 1 | 2022 | Role of Human-AI Interaction in Selective Prediction · AAAI 2022 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.6 | 1 | 2022 | Learning Optimal Conformal Classifiers · ICLR 2022 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.4 | 1 | 2020 | Adversarially Robust Representations with Smooth Encoders · ICLR 2020 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.4 | 1 | 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.4 | 1 | 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.4 | 1 | 2020 | The Autoencoding Variational Autoencoder · NeurIPS 2020 |
Machine learning › Generative modeling › generative adversarial network
StyleGAN |
0.4 | 1 | 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled Representations · CVPR 2020 |
Machine learning › Generative modeling
variational autoencoder |
0.4 | 1 | 2020 | The Autoencoding Variational Autoencoder · NeurIPS 2020 |
Machine learning › Optimization for machine learning
second-order optimization |
0.4 | 2 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 Stochastic Quasi-Newton Langevin Monte Carlo · ICML 2016 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
0.3 | 2 | 2018 | Stochastic Quasi-Newton Langevin Monte Carlo · ICML 2016 Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
stochastic gradient MCMC |
0.3 | 2 | 2018 | Stochastic Quasi-Newton Langevin Monte Carlo · ICML 2016 Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Machine learning › Optimization for machine learning › parallel optimization
asynchronous parallel optimization |
0.3 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Robotics › Robot navigation and mapping
localization |
0.3 | 1 | 2018 | EndoSensorFusion: Particle Filtering-Based Multi-Sensory Data Fusion with Switching State-Space Model for Endoscopic Capsule Robots · ICRA 2018 |
Robotics › Robot navigation and mapping › localization
multi-sensor localization |
0.3 | 1 | 2018 | EndoSensorFusion: Particle Filtering-Based Multi-Sensory Data Fusion with Switching State-Space Model for Endoscopic Capsule Robots · ICRA 2018 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.3 | 1 | 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex Optimization · ICML 2018 |
Audio and music processing
music information retrieval |
0.3 | 3 | 2015 | Inferring Metrical Structure in Music Using Particle Filters · IEEE ACM Trans. Audio Speech Lang. Process. 2015 Learning the beta-Divergence in Tweedie Compound Poisson Matrix Factorization Models · ICML (3) 2013 Tempo tracking and rhythm quantization by sequential Monte Carlo · NIPS 2001 |
Machine learning › Optimization for machine learning › second-order optimization
quasi-newton method |
0.2 | 1 | 2016 | Stochastic Quasi-Newton Langevin Monte Carlo · ICML 2016 |
Bioinformatics and computational biology › functional genomics
functional enrichment analysis |
0.2 | 1 | 2016 | CLUSTERnGO: a user-defined modelling platform for two-stage clustering of time-series data · Bioinform. 2016 |
Bioinformatics and computational biology › functional genomics › functional enrichment analysis
gene ontology analysis |
0.2 | 1 | 2016 | CLUSTERnGO: a user-defined modelling platform for two-stage clustering of time-series data · Bioinform. 2016 |
Methods — techniques the papers use, named apart from their topics
magnetic laplacian eigenvectors · 1.3directional random walk encoding · 1.3selective prediction · 1.1hypothesis testing · 0.8effect size analysis · 0.8human-in-the-loop experiments · 0.6human-in-the-loop experiment · 0.6distribution shift analysis · 0.6conformal prediction · 0.6benchmark evaluation · 0.6particle filter · 0.5adversarial mixing · 0.4sensor fusion · 0.3recurrent neural network · 0.3ergodic convergence analysis · 0.3entropy agglomeration · 0.3asynchronous parallelization · 0.3L-BFGS · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating medical AI systems in dermatology under uncertain ground truth
David Stutz, A. Taylan Cemgil, Abhijit Guha Roy, Tatiana Matejovicova, Melih Barsbey, Patricia Strachan, Mike Schaekermann, Jan Freyberg, Rajeev Rikhye, Beverly Freeman, Javier Perez Matos, Umesh Telang, Dale R. Webster, Gregory S. Corrado, Yossi Matias, Pushmeet Kohli, Yun Liu 0013, Arnaud Doucet, Alan Karthikesalingam |
Medical Image Anal. | 2 |
| 2024 | Evaluating Model Bias Requires Characterizing its MistakesabstractThe ability to properly benchmark model performance in the face of spurious correlations is important to both build better predictors and increase confidence that models are operating as intended. We demonstrate that characterizing (as opposed to simply quantifying) model mistakes across subgroups is pivotal to properly reflect model biases, which are ignored by standard metrics such as worst-group accuracy or accuracy gap. Inspired by the hypothesis testing framework, we introduce SkewSize, a principled and flexible metric that captures bias from mistakes in a model's predictions. It can be used in multi-class settings or generalised to the open vocabulary setting of generative models. SkewSize is an aggregation of the effect size of the interaction between two categorical variables: the spurious variable representing the bias attribute the model's prediction. We demonstrate the utility of SkewSize in multiple settings including: standard vision models trained on synthetic data, vision models trained on ImageNet, and large scale vision-and-language models from the BLIP-2 family. In each case, the proposed SkewSize is able to highlight biases not captured by other metrics, while also providing insights on the impact of recently proposed techniques, such as instruction tuning. Isabela Albuquerque, Jessica Schrouff, David Warde-Farley, A. Taylan Cemgil, Sven Gowal, Olivia Wiles |
ICML | 4 |
| 2023 | Transformers Meet Directed GraphsabstractTransformers were originally proposed as a sequence-to-sequence model for text but have become vital for a wide range of modalities, including images, audio, video, and undirected graphs. However, transformers for directed graphs are a surprisingly underexplored topic, despite their applicability to ubiquitous domains, including source code and logic circuits. In this work, we propose two direction- and structure-aware positional encodings for directed graphs: (1) the eigenvectors of the Magnetic Laplacian — a direction-aware generalization of the combinatorial Laplacian; (2) directional random walk encodings. Empirically, we show that the extra directionality information is useful in various downstream tasks, including correctness testing of sorting networks and source code understanding. Together with a data-flow-centric graph construction, our model outperforms the prior state of the art on the Open Graph Benchmark Code2 relatively by 14.7%. Simon Geisler, Yujia Li 0001, Daniel J. Mankowitz, A. Taylan Cemgil, Stephan Günnemann, Cosmin Paduraru |
ICML | 4 |
| 2023 | Probabilistic indoor tracking of Bluetooth Low-Energy beacons
F. Serhan Danis, Cem Ersoy, A. Taylan Cemgil |
Perform. Evaluation | 3 |
| 2022 | Role of Human-AI Interaction in Selective PredictionabstractRecent work has shown the potential benefit of selective prediction systems that can learn to defer to a human when the predictions of the AI are unreliable, particularly to improve the reliability of AI systems in high-stakes applications like healthcare or conservation. However, most prior work assumes that human behavior remains unchanged when they solve a prediction task as part of a human-AI team as opposed to by themselves. We show that this is not the case by performing experiments to quantify human-AI interaction in the context of selective prediction. In particular, we study the impact of communicating different types of information to humans about the AI system's decision to defer. Using real-world conservation data and a selective prediction system that improves expected accuracy over that of the human or AI system working individually, we show that this messaging has a significant impact on the accuracy of human judgements. Our results study two components of the messaging strategy: 1) Whether humans are informed about the prediction of the AI system and 2) Whether they are informed about the decision of the selective prediction system to defer. By manipulating these messaging components, we show that it is possible to significantly boost human performance by informing the human of the decision to defer, but not revealing the prediction of the AI. We therefore show that it is vital to consider how the decision to defer is communicated to a human when designing selective prediction systems, and that the composite accuracy of a human-AI team must be carefully evaluated using a human-in-the-loop framework. Elizabeth Bondi-Kelly, Raphael Koster, Hannah Sheahan, Martin J. Chadwick, Yoram Bachrach, A. Taylan Cemgil, Ulrich Paquet, Krishnamurthy Dvijotham |
AAAI | 6 |
| 2022 | Learning Optimal Conformal Classifiers
David Stutz, Krishnamurthy Dvijotham, A. Taylan Cemgil, Arnaud Doucet |
ICLR | 3 |
| 2022 | A Fine-Grained Analysis on Distribution Shift
Olivia Wiles, Sven Gowal, Florian Stimberg, Sylvestre-Alvise Rebuffi, Ira Ktena, Krishnamurthy Dvijotham, A. Taylan Cemgil |
ICLR | 7 |
| 2022 | Evaluating the Adversarial Robustness of Adaptive Test-time DefensesabstractAdaptive defenses, which optimize at test time, promise to improve adversarial robustness. We categorize such adaptive test-time defenses, explain their potential benefits and drawbacks, and evaluate a representative variety of the latest adaptive defenses for image classification. Unfortunately, none significantly improve upon static defenses when subjected to our careful case study evaluation. Some even weaken the underlying static model while simultaneously increasing inference computation. While these results are disappointing, we still believe that adaptive test-time defenses are a promising avenue of research and, as such, we provide recommendations for their thorough evaluation. We extend the checklist of Carlini et al. (2019) by providing concrete steps specific to adaptive defenses. Francesco Croce, Sven Gowal, Thomas Brunner, Evan Shelhamer, Matthias Hein 0001, A. Taylan Cemgil |
ICML | 6 |
| 2022 | Tracking a Mobile Beacon: A Purely Probabilistic ApproachabstractWe construct a practical and real-time probabilistic framework for fine target tracking. The practicality comes from the application of the forward algorithm and the small parameter set used to build the hidden Markov model (HMM). A Bluetooth Low-Energy (BLE) beacon navigating in the environment publishes BLE packets which are captured by the stationary sensors. Fingerprints are formed by collecting received signal strength indicators (RSSI) of these packets, which are then processed into a high resolution emission matrix using a histogram combination technique. We convert the map of the area into a grid structure, the resolution of which is controlled by the grid cell size. The transition matrices are built by Gaussian blur masks parametrized by the size and diffusion extent. As the transition matrix is highly sparse, we make the exact inference tractable by adopting a sparse matrix representation and by intelligently controlling the mask size, diffusion factor and grid cell size. Filtering can then be directly performed by the forward algorithm given a series of real RSSI measurements along real trajectories. We measure the performance of the system by comparing the most likely positions at each step with the ground truth positions. We achieve promising results and evaluate the approach also by the runtime and memory usage. F. Serhan Danis, Cem Ersoy, A. Taylan Cemgil |
MASCOTS | 3 |
| 2022 | Does your dermatology classifier know what it doesn't know? Detecting the long-tail of unseen conditions
Abhijit Guha Roy, Jie Ren 0006, Shekoofeh Azizi, Aaron Loh, Vivek Natarajan, Basil Mustafa, Nick Pawlowski, Jan Freyberg, Zachary Beaver, Nam Sy Vo, Peggy Bui, Samantha Winter, Patricia MacWilliams, Gregory S. Corrado, Umesh Telang, Yun Liu 0013, A. Taylan Cemgil, Alan Karthikesalingam, Balaji Lakshminarayanan, Jim Winkens |
Medical Image Anal. | 18 |
| 2022 | An indoor localization dataset and data collection framework with high precision position annotation
F. Serhan Danis, Ahmet Teoman Naskali, A. Taylan Cemgil, Cem Ersoy |
Pervasive Mob. Comput. | 3 |
| 2022 | Dirichlet-Luce choice model for learning from interactions
Gökhan Çapan, Ilker Gündogdu, Ali Caner Türkmen, A. Taylan Cemgil |
User Model. User Adapt. Interact. | 4 |
| 2021 | Unbiased gradient estimation for variational auto-encoders using coupled Markov chainsabstractThe variational auto-encoder (VAE) is a deep latent variable model that has two neural networks in an autoencoder-like architecture; one of them parameterizes the model’s likelihood. Fitting its parameters via maximum likelihood (ML) is challenging since the computation of the marginal likelihood involves an intractable integral over the latent space; thus the VAE is trained instead by maximizing a variational lower bound. Here, we develop a ML training scheme for VAEs by introducing unbiased estimators of the log-likelihood gradient. We obtain the estimators by augmenting the latent space with a set of importance samples, similarly to the importance weighted auto-encoder (IWAE), and then constructing a Markov chain Monte Carlo coupling procedure on this augmented space. We provide the conditions under which the estimators can be computed in finite time and with finite variance. We show experimentally that VAEs fitted with unbiased estimators exhibit better predictive performance. Francisco J. R. Ruiz, Michalis K. Titsias, A. Taylan Cemgil, Arnaud Doucet |
UAI | 3 |
| 2020 | Achieving Robustness in the Wild via Adversarial Mixing With Disentangled RepresentationsabstractRecent research has made the surprising finding that state-of-the-art deep learning models sometimes fail to generalize to small variations of the input. Adversarial training has been shown to be an effective approach to overcome this problem. However, its application has been limited to enforcing invariance to analytically defined transformations like lp-norm bounded perturbations. Such perturbations do not necessarily cover plausible real-world variations that preserve the semantics of the input (such as a change in lighting conditions). In this paper, we propose a novel approach to express and formalize robustness to these kinds of real-world transformations of the input. The two key ideas underlying our formulation are (1) leveraging disentangled representations of the input to define different factors of variations, and (2) generating new input images by adversarially composing the representations of different images. We use a StyleGAN model to demonstrate the efficacy of this framework. Specifically, we leverage the disentangled latent representations computed by a StyleGAN model to generate perturbations of an image that are similar to real-world variations (like adding make-up, or changing the skin-tone of a person) and train models to be invariant to these perturbations. Extensive experiments show that our method improves generalization and reduces the effect of spurious correlations (reducing the error rate of a "smile" detector by 21% for example). Sven Gowal, Chongli Qin, Po-Sen Huang, A. Taylan Cemgil, Krishnamurthy Dvijotham, Timothy A. Mann, Pushmeet Kohli |
CVPR | 4 |
| 2020 | Adversarially Robust Representations with Smooth Encoders
A. Taylan Cemgil, Sumedh Ghaisas, Krishnamurthy Dvijotham, Pushmeet Kohli |
ICLR | 1 |
| 2020 | The Autoencoding Variational AutoencoderabstractDoes a Variational AutoEncoder (VAE) consistently encode typical samples generated from its decoder? This paper shows that the perhaps surprising answer to this question is `No'; a (nominally trained) VAE does not necessarily amortize inference for typical samples that it is capable of generating. We study the implications of this behaviour on the learned representations and also the consequences of fixing it by introducing a notion of self consistency. Our approach hinges on an alternative construction of the variational approximation distribution to the true posterior of an extended VAE model with a Markov chain alternating between the encoder and the decoder. The method can be used to train a VAE model from scratch or given an already trained VAE, it can be run as a post processing step in an entirely self supervised way without access to the original training data. Our experimental analysis reveals that encoders trained with our self-consistency approach lead to representations that are robust (insensitive) to perturbations in the input introduced by adversarial attacks. We provide experimental results on the ColorMnist and CelebA benchmark datasets that quantify the properties of the learned representations and compare the approach with a baseline that is specifically trained for the desired property. A. Taylan Cemgil, Sumedh Ghaisas, Krishnamurthy Dvijotham, Sven Gowal, Pushmeet Kohli |
NeurIPS | 1 |
| 2020 | Clustering Event Streams With Low Rank Hawkes ProcessesabstractWe introduce a fast algorithm for parameter estimation in multidimensional Hawkes processes, a widely used class of temporal point processes for mutually exciting discrete event data. Our approach assumes a low-rank structure on the infectivity parameter of the multidimensional Hawkes process, and relies on a method of moments estimator. Notably, it requires only a single scan of the data, and consistently recovers an accurate representation of the underlying graph structure, while sidestepping numerical stability issues inherent in Hawkes process estimation. Finally, we make connections between our method and spectral clustering, and observe that our contributions result in natural methods for clustering temporal point processes. Our algorithm can be used for community detection and graph cluster discovery in large networks of asynchronous event streams such as high-dimensional neural spike trains, log streams of large computer networks, or high-frequency financial data. We present favorable empirical results on synthetic data, and an application to clustering currency pairs via high-frequency price jumps. Ali Caner Türkmen, Gökhan Çapan, A. Taylan Cemgil |
IEEE Signal Process. Lett. | 3 |
| 2020 | Data Sharing via Differentially Private Coupled Matrix FactorizationabstractWe address the privacy-preserving data-sharing problem in a distributed multiparty setting. In this setting, each data site owns a distinct part of a dataset and the aim is to estimate the parameters of a statistical model conditioned on the complete data without any site revealing any information about the individuals in their own parts. The sites want to maximize the utility of the collective data analysis while providing privacy guarantees for theirown portion of the dataas well as foreach participating individual. Our first contribution is to classify these different privacy requirements as (i)site-leveland (ii)user-leveldifferential privacy and present formal privacy guarantees for these two cases under the model of differential privacy. To satisfy a stronger form of differential privacy, we use a variant of differential privacy which islocal differential privacywhere the sensitive data is perturbed with a randomized response mechanism prior to the estimation. In this study, we assume that the data instances that are partitioned between several parties are arranged as matrices. A natural statistical model for this distributed scenario is coupled matrix factorization. We present two generic frameworks for privatizing Bayesian inference for coupled matrix factorization models that are able to guarantee proposed differential privacy notions based on the privacy requirements of the model. To privatize Bayesian inference, we first exploit the connection between differential privacy and sampling from a Bayesian posterior via stochastic gradient Langevin dynamics and then derive an efficient coupled matrix factorization method. In the local privacy context, we propose two models that have an additional privatization mechanism to achieve a stronger measure of privacy and introduce a Gibbs sampling based algorithm. We demonstrate that the proposed methods are able to provide good prediction accuracy on synthetic and real datasets while adhering to the introduced privacy constraints. Beyza Ermis, A. Taylan Cemgil |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | Audio Source Separation Using Variational Autoencoders and Weak Class SupervisionabstractIn this letter, we propose a source separation method that is trained by observing the mixtures and the class labels of the sources present in the mixture without any access to isolated sources. Since our method does not require source class labels for every time-frequency bin but only a single label for each source constituting the mixture signal, we call this scenario as weak class supervision. We associate a variational autoencoder (VAE) with each source class within a nonnegative (compositional) model. Each VAE provides a prior model to identify the signal from its associated class in a sound mixture. After training the model on mixtures, we obtain a generative model for each source class and demonstrate our method on one-second mixtures of utterances of digits from 0 to 9. We show that the separation performance obtained by source class supervision is as good as the performance obtained by source signal supervision. Ertug Karamatli, A. Taylan Cemgil, Serap Kirbiz |
IEEE Signal Process. Lett. | 2 |
| 2019 | Estimating Network Flow Length Distributions via Bayesian Nonnegative Tensor FactorizationabstractIn this paper, we develop a framework to estimate network flow length distributions in terms of the number of packets. We model the network flow length data as a three-way array with day-of-week, hour-of-day, and flow length as entities where we observe a count. In a high-speed network, only a sampled version of such an array can be observed and reconstructing the true flow statistics from fewer observations becomes a computational problem. We formulate the sampling process as matrix multiplication so that any sampling method can be used in our framework as long as its sampling probabilities are written in matrix form. We demonstrate our framework on a high-volume real-world data set collected from a mobile network provider with a random packet sampling and a flow-based packet sampling methods. We show that modeling the network data as a tensor improves estimations of the true flow length histogram in both sampling methods. Baris Kurt, A. Taylan Cemgil, Gunes Karabulut-Kurt, Engin Zeydan |
Wirel. Commun. Mob. Comput. | 2 |
| 2018 | Asynchronous Stochastic Quasi-Newton MCMC for Non-Convex OptimizationabstractRecent studies have illustrated that stochastic gradient Markov Chain Monte Carlo techniques have a strong potential in non-convex optimization, where local and global convergence guarantees can be shown under certain conditions. By building up on this recent theory, in this study, we develop an asynchronous-parallel stochastic L-BFGS algorithm for non-convex optimization. The proposed algorithm is suitable for both distributed and shared-memory settings. We provide formal theoretical analysis and show that the proposed method achieves an ergodic convergence rate of ${\cal O}(1/\sqrt{N})$ ($N$ being the total number of iterations) and it can achieve a linear speedup under certain conditions. We perform several experiments on both synthetic and real datasets. The results support our theory and show that the proposed algorithm provides a significant speedup over the recently proposed synchronous distributed L-BFGS algorithm. Umut Simsekli, Çagatay Yildiz, Thanh Huy Nguyen 0001, A. Taylan Cemgil, Gaël Richard |
ICML | 4 |
| 2018 | EndoSensorFusion: Particle Filtering-Based Multi-Sensory Data Fusion with Switching State-Space Model for Endoscopic Capsule RobotsabstractA reliable, real time, multi-sensor fusion functionality is crucial for localization of actively controlled capsule endoscopy robots, which are an emerging, minimally invasive diagnostic and therapeutic technology for the gastrointestinal (GI) tract. In this study, we propose a novel multi-sensor fusion approach based on a particle filter that incorporates an online estimation of sensor reliability and a non-linear kinematic model learned by a recurrent neural network. Our method sequentially estimates the true robot pose from noisy pose observations delivered by multiple sensors. We experimentally test the method using 5 degree-of-freedom (5-DoF) absolute pose measurement by a magnetic localization system and a 6-DoF relative pose measurement by visual odometry. In addition, the proposed method is capable of detecting and handling sensor failures by ignoring corrupted data, providing the robustness expected of a medical device. Detailed analyses and evaluations are presented using ex vivo experiments on a porcine stomach model, proving that our system achieves high translational and rotational accuracies for different types of endoscopic capsule robot trajectories. Mehmet Turan, Yasin Almalioglu, Hunter B. Gilbert, Helder Araújo, A. Taylan Cemgil, Metin Sitti |
ICRA | 5 |
| 2018 | An intelligent cyber security system against DDoS attacks in SIP networks
Murat Semerci, A. Taylan Cemgil, Bülent Sankur |
Comput. Networks | 2 |
| 2018 | Efficient Bayesian Model Selection in PARAFAC via Stochastic Thermodynamic IntegrationabstractParallel factor analysis (PARAFAC) is one of the most popular tensor factorization models. Even though it has proven successful in diverse application fields, the performance of PARAFAC usually hinges up on the rank of the factorization, which is typically specified manually by the practitioner. In this study, we develop a novel parallel and distributed Bayesian model selection technique for rank estimation in large-scale PARAFAC models. The proposed approach integrates ideas from the emerging field of stochastic gradient Markov Chain Monte Carlo, statistical physics, and distributed stochastic optimization. As opposed to the existing methods, which are based on some heuristics, our method has a clear mathematical interpretation, and has significantly lower computational requirements, thanks to data subsampling and parallelization. We provide formal theoretical analysis on the bias induced by the proposed approach. Our experiments on synthetic and large-scale real datasets show that our method is able to find the optimal model order while being significantly faster than the state-of-the-art. Thanh Huy Nguyen 0001, Umut Simsekli, Gaël Richard, A. Taylan Cemgil |
IEEE Signal Process. Lett. | 4 |
| 2017 | Parallelized Stochastic Gradient Markov Chain Monte Carlo algorithms for non-negative matrix factorizationabstractStochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have become popular in modern data analysis problems due to their computational efficiency. Even though they have proved useful for many statistical models, the application of SG-MCMC to non-negative matrix factorization (NMF) models has not yet been extensively explored. In this study, we develop two parallel SG-MCMC algorithms for a broad range of NMF models. We exploit the conditional independence structure of the NMF models and utilize a stratified sub-sampling approach for enabling parallelization. We illustrate the proposed algorithms on an image restoration task and report encouraging results. Umut Simsekli, Alain Durmus, Roland Badeau, Gaël Richard, Eric Moulines, A. Taylan Cemgil |
ICASSP | 6 |
| 2016 | Stochastic thermodynamic integration: Efficient Bayesian model selection via stochastic gradient MCMCabstractModel selection is a central topic in Bayesian machine learning, which requires the estimation of the marginal likelihood of the data under the models to be compared. During the last decade, conventional model selection methods have lost their charm as they have high computational requirements. In this study, we propose a computationally efficient model selection method by integrating ideas from Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) literature and statistical physics. As opposed to conventional methods, the proposed method has very low computational needs and can be implemented almost without modifying existing SG-MCMC code. We provide an upper-bound for the bias of the proposed method. Our experiments show that, our method is 40 times as fast as the baseline method on finding the optimal model order in a matrix factorization problem. Umut Simsekli, Roland Badeau, Gaël Richard, A. Taylan Cemgil |
ICASSP | 4 |
| 2016 | A generalized Bayesian model for tracking long metrical cycles in acoustic music signalsabstractMost musical phenomena involve repetitive structures that enable listeners to track meter, i.e. the tactus or beat, the longer over-arching measure or bar, and possibly other related layers. Meters with long measure duration, sometimes lasting more than a minute, occur in many music cultures, e.g. from India, Turkey, and Korea. However, current meter tracking algorithms, which were devised for cycles of a few seconds length, cannot process such structures accurately. We present a novel generalization to an existing Bayesian model for meter tracking that overcomes this limitation. The proposed model is evaluated on a set of Indian Hindustani music recordings, and we document significant performance increase over the previous models. The presented model opens the way for computational analysis of performances with long metrical cycles, and has important applications in music studies as well as in commercial applications that involve such musics. Ajay Srinivasamurthy, Andre Holzapfel, A. Taylan Cemgil, Xavier Serra |
ICASSP | 3 |
| 2016 | Stochastic Quasi-Newton Langevin Monte CarloabstractRecently, Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have been proposed for scaling up Monte Carlo computations to large data problems. Whilst these approaches have proven useful in many applications, vanilla SG-MCMC might suffer from poor mixing rates when random variables exhibit strong couplings under the target densities or big scale differences. In this study, we propose a novel SG-MCMC method that takes the local geometry into account by using ideas from Quasi-Newton optimization methods. These second order methods directly approximate the inverse Hessian by using a limited history of samples and their gradients. Our method uses dense approximations of the inverse Hessian while keeping the time and memory complexities linear with the dimension of the problem. We provide a formal theoretical analysis where we show that the proposed method is asymptotically unbiased and consistent with the posterior expectations. We illustrate the effectiveness of the approach on both synthetic and real datasets. Our experiments on two challenging applications show that our method achieves fast convergence rates similar to Riemannian approaches while at the same time having low computational requirements similar to diagonal preconditioning approaches. Umut Simsekli, Roland Badeau, A. Taylan Cemgil, Gaël Richard |
ICML | 3 |
| 2016 | A Network Monitoring System for High Speed Network TrafficabstractMonitoring network statistics is important for the maintenance and infrastructure planning for the network service providers. In this demonstration, we will showcase an initial analysis of a general purpose network monitoring platform for high speed mobile networks. The developed platform is the basis for performing complex real-time analysis such as application usage behaviour, security analysis, infrastructure planning. We have used the platform for real-time flow size and length monitoring with packet sampling. Baris Kurt, Engin Zeydan, Utku Yabas, Ilyas Alper Karatepe, Gunes Karabulut-Kurt, A. Taylan Cemgil |
SECON | 6 |
| 2016 | CLUSTERnGO: a user-defined modelling platform for two-stage clustering of time-series dataabstractMOTIVATION: Simple bioinformatic tools are frequently used to analyse time-series datasets regardless of their ability to deal with transient phenomena, limiting the meaningful information that may be extracted from them. This situation requires the development and exploitation of tailor-made, easy-to-use and flexible tools designed specifically for the analysis of time-series datasets. RESULTS: We present a novel statistical application called CLUSTERnGO, which uses a model-based clustering algorithm that fulfils this need. This algorithm involves two components of operation. Component 1 constructs a Bayesian non-parametric model (Infinite Mixture of Piecewise Linear Sequences) and Component 2, which applies a novel clustering methodology (Two-Stage Clustering). The software can also assign biological meaning to the identified clusters using an appropriate ontology. It applies multiple hypothesis testing to report the significance of these enrichments. The algorithm has a four-phase pipeline. The application can be executed using either command-line tools or a user-friendly Graphical User Interface. The latter has been developed to address the needs of both specialist and non-specialist users. We use three diverse test cases to demonstrate the flexibility of the proposed strategy. In all cases, CLUSTERnGO not only outperformed existing algorithms in assigning unique GO term enrichments to the identified clusters, but also revealed novel insights regarding the biological systems examined, which were not uncovered in the original publications. AVAILABILITY AND IMPLEMENTATION: The C++ and QT source codes, the GUI applications for Windows, OS X and Linux operating systems and user manual are freely available for download under the GNU GPL v3 license at http://www.cmpe.boun.edu.tr/content/CnG. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Isik Baris Fidaner, Ayca Cankorur-Cetinkaya, Duygu Dikicioglu, Betül Kirdar, A. Taylan Cemgil, Stephen G. Oliver |
Bioinform. | 5 |
| 2015 | Section-level modeling of musical audio for linking performances to scores in Turkish makam musicabstractSection linking aims at relating structural units in the notation of a piece of music to their occurrences in a performance of the piece. In this paper, we address this task by presenting a score-informed hierarchical Hidden Markov Model (HHMM) for modeling musical audio signals on the temporal level of sections present in a composition, where the main idea is to explicitly model the long range and hierarchical structure of music signals. So far, approaches based on HHMM or similar methods were mainly developed for a note-to-note alignment, i.e. an alignment based on shorter temporal units than sections. Such approaches, however, are conceptually problematic when the performances differ substantially from the reference score due to interpretation and improvisation, a very common phenomenon, for instance, in Turkish makam music. In addition to having low computational complexity compared to note-to-note alignment and achieving a transparent and elegant model, the experimental results show that our method outperforms a previously presented approach on a Turkish makam music corpus. Andre Holzapfel, Umut Simsekli, Sertan Sentürk, A. Taylan Cemgil |
ICASSP | 4 |
| 2015 | Learning mixed divergences in coupled matrix and tensor factorization modelsabstractCoupled tensor factorization methods are useful for sensor fusion, combining information from several related datasets by simultaneously approximating them by products of latent tensors. In these methods, the choice of a suitable optimization criteria becomes difficult as observed datasets may have different statistical characteristics and their relative importance for the task at hand can vary. In this paper, we present an algorithmic framework for coupled factorization that, while estimating a latent factorization also estimates a specific ß-divergence for each dataset as well as the relative weights in an overall additive cost function. We evaluate the proposed method on both synthetical and real datasets, where we apply our methods on a link prediction problem. The results show that our method outperforms the state-of-the-art by a significant margin. Umut Simsekli, A. Taylan Cemgil, Beyza Ermis |
ICASSP | 2 |
| 2015 | Link prediction in heterogeneous data via generalized coupled tensor factorization
Beyza Ermis, Evrim Acar, A. Taylan Cemgil |
Data Min. Knowl. Discov. | 3 |
| 2015 | Alpha-Stable Matrix FactorizationabstractMatrix factorization (MF) models have been widely used in data analysis. Even though they have been shown to be useful in many applications, classical MF models often fall short when the observed data are impulsive and contain outliers. In this study, we present$\alpha $MF, a MF model with$\alpha $-stable observations. Stable distributions are a family of heavy-tailed distributions that is particularly suited for such impulsive data. We develop a Markov Chain Monte Carlo method, namely a Gibbs sampler, for making inference in the model. We evaluate our model on both synthetic and real audio applications. Our experiments on speech enhancement show that$\alpha $MF yields superior performance to a popular audio processing model in terms of objective measures. Furthermore,$\alpha $MF provides a theoretically sound justification for recent empirical results obtained in audio processing. Umut Simsekli, Antoine Liutkus, A. Taylan Cemgil |
IEEE Signal Process. Lett. | 3 |
| 2015 | A Probabilistic Model-Based Approach for Aligning Multiple Audio SequencesabstractWe formulate the alignment problem of multiple and partially overlapping audio sequences in a probabilistic framework. We define and compare five generative models for several time varying features extracted from audio clips that are recorded independently and asynchronously. For each model, we derive the associated scoring function that evaluates the quality of an alignment. The matching is then achieved with a sequential algorithm. The derived score functions are also able to identify the cases where the sequences do not overlap and handle multiple sequences where no sequence is covering the entire timeline. The simulation results on real data suggest that the approach is able to handle difficult, ambiguous scenarios and partial matchings where simple baseline methods such as correlation fail. Dogac Basaran, A. Taylan Cemgil, Emin Anarim |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Inferring Metrical Structure in Music Using Particle FiltersabstractIn this paper, we propose a new state-of-the-art particle filter (PF) system to infer the metrical structure of musical audio signals. The new inference method is designed to overcome the problem of PFs in multi-modal probability distributions, which arise due to tempo and phase ambiguities in musical rhythm representations. We compare the new method with a hidden Markov model (HMM) system and several other PF schemes in terms of performance, speed and scalability on several audio datasets. We demonstrate that using the proposed system the computational complexity can be reduced drastically in comparison to the HMM while maintaining the same order of beat tracking accuracy. Therefore, for the first time, the proposed system allows fast meter inference in a high-dimensional state space, spanned by the three components of tempo, type of rhythm, and position in a metric cycle. Florian Krebs, Andre Holzapfel, A. Taylan Cemgil, Gerhard Widmer |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | SMC samplers for multiresolution audio sequence alignmentabstractIn our previous work, we formulated multiple audio sequence alignment in a probabilistic framework [1]. Here, we extend the model for multi resolution alignment and focus on pairwise cases. We defined a similarity based approach for binary feature sequences and integrate it into a new generative model. We modify themodel formulti resolution case and the matching is achieved with a SequentialMonte Carlo Sampler (SMCS) which uses low resolution models as bridge distributions. The simulation results on real data sets suggest that our method is very robust and efficient under very noisy conditions with proper choices of model parameters. Dogac Basaran, A. Taylan Cemgil, Emin Anarim |
ICASSP | 2 |
| 2013 | Learning the beta-Divergence in Tweedie Compound Poisson Matrix Factorization ModelsabstractIn this study, we derive algorithms for estimating mixed β-divergences. Such cost functions are useful for Nonnegative Matrix and Tensor Factorization models with a compound Poisson observation model. Compound Poisson is a particular Tweedie model, an important special case of exponential dispersion models characterized by the fact that the variance is proportional to a power function of the mean. There are several well known matrix and tensor factorization algorithms that minimize the β-divergence; these estimate the mean parameter. The probabilistic interpretation gives us more flexibility and robustness by providing us additional tunable parameters such as power and dispersion. Estimation of the power parameter is useful for choosing a suitable divergence and estimation of dispersion is useful for data driven regularization and weighting in collective/coupled factorization of heterogeneous datasets. We present three inference algorithms for both estimating the factors and the additional parameters of the compound Poisson distribution. The methods are evaluated on two applications: modeling symbolic representations for polyphonic music and lyric prediction from audio features. Our conclusion is that the compound poisson based factorization models can be useful for sparse positive data. Umut Simsekli, A. Taylan Cemgil, Yusuf Kenan Yilmaz |
ICML (3) | 2 |
| 2013 | Summary Statistics for Partitionings and Feature AllocationsabstractInfinite mixture models are commonly used for clustering. One can sample from the posterior of mixture assignments by Monte Carlo methods or find its maximum a posteriori solution by optimization. However, in some problems the posterior is diffuse and it is hard to interpret the sampled partitionings. In this paper, we introduce novel statistics based on block sizes for representing sample sets of partitionings and feature allocations. We develop an element-based definition of entropy to quantify segmentation among their elements. Then we propose a simple algorithm called entropy agglomeration (EA) to summarize and visualize this information. Experiments on various infinite mixture posteriors as well as a feature allocation dataset demonstrate that the proposed statistics are useful in practice. Isik Baris Fidaner, A. Taylan Cemgil |
NIPS | 2 |
| 2013 | Single-Channel Speech-Music Separation for Robust ASR With Mixture ModelsabstractIn this study, we describe a mixture model based single-channel speech-music separation method. Given a catalog of background music material, we propose a generative model for the superposed speech and music spectrograms. The background music signal is assumed to be generated by a jingle in the catalog. The background music component is modeled by a scaled conditional mixture model representing the jingle. The speech signal is modeled by a probabilistic model, which is similar to the probabilistic interpretation of Non-negative Matrix Factorization (NMF) model. The parameters of the speech model is estimated in a semi-supervised manner from the mixed signal. The approach is tested with Poisson and complex Gaussian observation models that correspond respectively to Kullback-Leibler (KL) and Itakura-Saito (IS) divergence measures. Our experiments show that the proposed mixture model outperforms a standard NMF method both in speech-music separation and automatic speech recognition (ASR) tasks. These results are further improved using Markovian prior structures for temporal continuity between the jingle frames. Our test results with real data show that our method increases the speech recognition performance. Cemil Demir, Murat Saraclar, A. Taylan Cemgil |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Score guided audio restoration via generalised coupled tensor factorisationabstractGeneralised coupled tensor factorisation is a recently proposed algorithmic framework for simultaneously estimating tensor factorisation models where several observed tensors can share a set of latent factors. This paper proposes a model in this framework for coupled factorisation of piano spectrograms and piano roll representations to solve audio interpolation and restoration problem. The model incorporates temporal and harmonic information from an approximate musical score (not necessarily belonging to the played piece), and spectral information from isolated piano sounds. The performance of the proposed approach is evaluated on the restoration of classical music pieces where we get about 5dB SNR improvement when 50% of data frames are missing. Umut Simsekli, Yusuf Kenan Yilmaz, A. Taylan Cemgil |
ICASSP | 3 |
| 2012 | Effect of speech priors in single-channel speech-music separation for ASR
Cemil Demir, A. Taylan Cemgil, Murat Saraclar |
INTERSPEECH | 2 |
| 2012 | Algorithms for probabilistic latent tensor factorization
Yusuf Kenan Yilmaz, A. Taylan Cemgil |
Signal Process. | 2 |
| 2011 | Gain estimation approaches in catalog-based single-channel speech-music separationabstractIn this study, we analyze the gain estimation problem of the catalog-based single-channel speech-music separation method, which we proposed previously. In the proposed method, assuming that we know a catalog of the background music, we developed a generative model for the superposed speech and music spectrograms. We represent the speech spectrogram by a Non-Negative Matrix Factorization (NMF) model and the music spectrogram by a conditional Poisson Mixture Model (PMM). In this model, we assume that the background music is generated by repeating and changing the gain of the jingle in the music catalog. Although the separation performance of the proposed method is satisfactory with known gain values, the performance decreases when the gain value of the jingle is unknown and has to be estimated. In this paper, we address the gain estimation problem of the catalog-based method and propose three different approaches to overcome this problem. One of these approaches is to use Gamma Markov Chain (GMC) probabilistic structure to impose the correlation between the gain parameters across the time frames. By using GMC, the gain parameter is estimated more accurately. The other approaches are maximum a posteriori (MAP) and piece-wise constant estimation (PCE) of the gain values. Although all three methods improve the separation performance as compared to the original method itself, GMC approach achieved the best performance. Cemil Demir, A. Taylan Cemgil, Murat Saraclar |
ASRU | 2 |
| 2011 | Semi-Supervised Single-Channel Speech-Music Separation for Automatic Speech RecognitionabstractIn this study, we propose a semi-supervised speech-music separation method which uses the speech, music and speech-music segments in a given segmented audio signal to separate speech and music signals from each other in the mixed speech-music segments. In this strategy, we assume, the background music of the mixed signal is partially composed of the repetition of the music segment in the audio. Therefore, we used a mixture model to represent the music signal. The speech signal is modeled using Non-negative Matrix Factorization (NMF) model. The prior model of the template matrix of the NMF model is estimated using the speech segment and updated using the mixed segment of the audio. The separation performance of the proposed method is evaluated in automatic speech recognition task. Cemil Demir, A. Taylan Cemgil, Murat Saraclar |
INTERSPEECH | 2 |
| 2011 | Generalised Coupled Tensor FactorisationabstractWe derive algorithms for generalised tensor factorisation (GTF) by building upon the well-established theory of Generalised Linear Models. Our algorithms are general in the sense that we can compute arbitrary factorisations in a message passing framework, derived for a broad class of exponential family distributions including special cases such as Tweedie's distributions corresponding to $\beta$-divergences. By bounding the step size of the Fisher Scoring iteration of the GLM, we obtain general updates for real data and multiplicative updates for non-negative data. The GTF framework is, then extended easily to address the problems when multiple observed tensors are factorised simultaneously. We illustrate our coupled factorisation approach on synthetic data as well as on a musical audio restoration problem. Yusuf Kenan Yilmaz, A. Taylan Cemgil, Umut Simsekli |
NIPS | 2 |
| 2011 | Annealed SMC Samplers for Nonparametric Bayesian Mixture ModelsabstractWe develop a novel online algorithm for posterior inference in Dirichlet Process Mixtures (DPM). Our method is based on the Sequential Monte Carlo (SMC) samplers framework that generalizes sequential importance sampling approaches. Unlike the existing methods, the framework enables us to retrospectively update long trajectories in the light of recent observations and this leads to sophisticated clustering update schemes and annealing strategies that seem to prevent the algorithm to get stuck around a local mode. The performance has been evaluated on a Bayesian Gaussian density estimation problem with an unknown number of mixture components. Our simulations suggest that the proposed annealing strategy outperforms conventional samplers. It also provides significantly smaller Monte Carlo standard error with respect to particle filtering given comparable computational resources. Yener Ülker, Bilge Günsel, A. Taylan Cemgil |
IEEE Signal Process. Lett. | 3 |
| 2011 | Bayesian Interpolation and Parameter Estimation in a Dynamic Sinusoidal ModelabstractIn this paper, we propose a method for restoring the missing or corrupted observations of nonstationary sinusoidal signals which are often encountered in music and speech applications. To model nonstationary signals, we use a time-varying sinusoidal model which is obtained by extending the static sinusoidal model into a dynamic sinusoidal model. In this model, the in-phase and quadrature components of the sinusoids are modeled as first-order Gauss-Markov processes. The inference scheme for the model parameters and missing observations is formulated in a Bayesian framework and is based on a Markov chain Monte Carlo method known as Gibbs sampler. We focus on the parameter estimation in the dynamic sinusoidal model since this constitutes the core of model-based interpolation. In the simulations, we first investigate the applicability of the model and then demonstrate the inference scheme by applying it to the restoration of lost audio packets on a packet-based network. The results show that the proposed method is a reasonable inference scheme for estimating unknown signal parameters and interpolating gaps consisting of missing/corrupted signal segments. Jesper Kjær Nielsen, Mads Græsbøll Christensen, A. Taylan Cemgil, Simon J. Godsill, Søren Holdt Jensen |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Bayesian Inference for Nonnegative Matrix Factor Deconvolution ModelsabstractIn this paper we develop a probabilistic interpretation and a full Bayesian inference for non-negative matrix deconvolution (NMFD) model. Our ultimate goal is unsupervised extraction of multiple sound objects from a single channel auditory scene. The proposed method facilitates automatic model selection and determination of the sparsity criteria. Our approach retains attractive features of standard NMFD based methods such as fast convergence and easy implementation. We demonstrate the use of this algorithm in the log-frequency magnitude spectrum domain, where we employ it to perform model order selection and control sparseness directly. Serap Kirbiz, A. Taylan Cemgil, Bilge Günsel |
ICPR | 2 |
| 2010 | Annealed SMC Samplers for Dirichlet Process Mixture ModelsabstractIn this work we propose a novel algorithm that approximates sequentially the Dirichlet Process Mixtures (DPM) model posterior. The proposed method takes advantage of the Sequential Monte Carlo (SMC) samplers framework to design an effective annealing procedure that prevents the algorithm to get trapped in a local mode. We evaluate the performance in a Bayesian density estimation problem with unknown number of components. The simulation results suggest that the proposed algorithm represents the target posterior much more accurately and provides significantly smaller Monte Carlo error when compared to particle filtering. Yener Ülker, Bilge Günsel, A. Taylan Cemgil |
ICPR | 3 |
| 2010 | Catalog-based single-channel speech-music separation
Cemil Demir, A. Taylan Cemgil, Murat Saraclar |
INTERSPEECH | 2 |
| 2010 | Gamma Markov Random Fields for Audio Source ModelingabstractIn many audio processing tasks, such as source separation, denoising or compression, it is crucial to construct realistic and flexible models to capture the physical properties of audio signals. This can be accomplished in the Bayesian framework through the use of appropriate prior distributions. In this paper, we describe a class of prior models called Gamma Markov random fields (GMRFs) to model the sparsity and the local dependency of the energies (i.e., variances) of time-frequency expansion coefficients. A GMRF model describes a non-normalised joint distribution over unobserved variance variables, where given the field the actual source coefficients are independent. Our construction ensures a positive coupling between the variance variables, so that signal energy changes smoothly over both axes to capture the temporal and spectral continuity. The coupling strength is controlled by a set of hyperparameters. Inference on the overall model is convenient because of the conditional conjugacy of all of the variables in the model, but automatic optimization of hyperparameters is crucial to obtain better fits. The marginal likelihood of the model is not available because of the intractable normalizing constant of GMRFs. In this paper, we optimize the hyperparameters of our GMRF-based audio model using contrastive divergence and compare this method to alternatives such as score matching and pseudolikelihood maximization where applicable. We present the performance of the GMRF models in denoising and single-channel source separation problems in completely blind scenarios, where all the hyperparameters are jointly estimated given only audio data. Onur Dikmen, A. Taylan Cemgil |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Generative Spectrogram Factorization Models for Polyphonic Piano TranscriptionabstractWe introduce a framework for probabilistic generative models of time–frequency coefficients of audio signals, using a matrix factorization parametrization to jointly model spectral characteristics such as harmonicity and temporal activations and excitations. The models represent the observed data as the superposition of statistically independent sources, and we consider variance-based models used in source separation and intensity-based models for non-negative matrix factorization. We derive a generalized expectation-maximization algorithm for inferring the parameters of the model and then adapt this algorithm for the task of polyphonic transcription of music using labeled training data. The performance of the system is compared to that of existing discriminative and model-based approaches on a dataset of solo piano music. Paul H. Peeling, A. Taylan Cemgil, Simon J. Godsill |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | A Bayesian Deconvolution Approach for Receiver Function AnalysisabstractIn this paper, we propose a Bayesian methodology for receiver function analysis, a key tool in determining the deep structure of the Earth's crust. We exploit the assumption of sparsity for receiver functions to develop a Bayesian deconvolution method as an alternative to the widely used iterative deconvolution. We model samples of a sparse signal as i.i.d. Student-t random variables. Gibbs sampling and variational Bayes techniques are investigated for our specific posterior inference problem. We used those techniques within the expectation-maximization (EM) algorithm to estimate our unknown model parameters. The superiority of the Bayesian deconvolution is demonstrated by the experiments on both simulated and real earthquake data. Sinan Yildirim, A. Taylan Cemgil, Mustafa Aktar, Yaman Ozakin, Aysin Ertüzün |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2009 | A hybrid method for deconvolution of Bernoulli-Gaussian processesabstractWe investigate a hybrid method which improves the quality of state inference and parameter estimation in blind deconvolution of a sparse source modeled by a Bernoulli-Gaussian process. In this problem, when both the signal and the filter are jointly estimated, the true posterior is typically highly multimodal. Therefore, when not properly initialized, standard stochastic inference methods, (MCEM, SEM or SAEM), tend to get stuck and suffer from poor convergence. In our approach, we first relax the Bernoulli-Gaussian prior model by a Student-t model. Our simulations suggest that deterministic inference in the relaxed model is not only efficient, but also provides a very good initialization for the Bernoulli-Gaussian model. We provide simulation studies that compare the results obtained with and without our initialization method for several combinations of state inference and parameter estimation methods used for the Bernoulli-Gaussian model. Sinan Yildirim, A. Taylan Cemgil, Aysin Ertüzün |
ICASSP | 2 |
| 2009 | Approximate Bayesian methods for kernel-based object tracking
Zoran Zivkovic, A. Taylan Cemgil, Ben J. A. Kröse |
Comput. Vis. Image Underst. | 2 |
| 2008 | Bayesian extensions to non-negative matrix factorisation for audio signal modellingabstractWe describe the underlying probabilistic generative signal model of non-negative matrix factorisation (NMF) and propose a realistic conjugate priors on the matrices to be estimated. A conjugate Gamma chain prior enables modelling the spectral smoothness of natural sounds in general, and other prior knowledge about the spectra of the sounds can be used without resorting to too restrictive techniques where some of the parameters are fixed. The resulting algorithm, while retaining the attractive features of standard NMF such as fast convergence and easy implementation, outperforms existing NMF strategies in a single channel audio source separation and detection task. Tuomas Virtanen, A. Taylan Cemgil, Simon J. Godsill |
ICASSP | 2 |
| 2007 | Strategies for Sequential Inference in Factorial Switching State Space ModelsabstractFactorial switching state space models are large hybrid time series models in which inference is intractable even in a single time slice. For the conditional Gaussian case, we derive a message propagation algorithm (upward-downward) that exploits the factorial structure of the model and facilitates computing messages without the need for inverting large matrices. Using the propagation algorithm as a sub-routine, we develop a Rao-Blackwellized Gibbs sampler and a variational approximation of structured mean field type to compute an approximate proposal density. These proposal are useful for both filtering or for marginal maximum a-posteriori estimates. We illustrate the utility of our approach on a large factorial state space model for polyphonic music transcription. A. Taylan Cemgil |
ICASSP (2) | 1 |
| 2007 | Sequential Inference of Rhythmic Structure in Musical AudioabstractThis paper presents a framework for the modelling of temporal characteristics of musical signals and an approximate, sequential Monte Carlo inference scheme which yields estimates of tempo and rhythmic pattern from onset-time data. These two features are quantified through the construction of a probabilistic dynamical model of a hidden 'bar-pointer' and a Poisson observation model. The capabilities of the system are demonstrated by tracking the tempo of a 2 against 3 polyrhythm and detecting a switch in rhythm in a MIDI performance. Nick Whiteley, A. Taylan Cemgil, Simon J. Godsill |
ICASSP (4) | 2 |
| 2007 | Bayesian methods for multimedia signal processingabstractIn the last years, there have been a significant growth of multimedia information processing applications that employ ideas from statistical machine learning and probabilistic modeling. In this paradigm, multimedia data (music, audio, video, images, text, ...) are viewed as realizations from highly structured stochastic processes. Once a model is constructed, several interesting problems such as transcription, coding, classification, restoration, tracking, source separation or resynthesis etc. can be formulated as Bayesian inference problems. In this context, graphical models provide a "language" to construct models for quantification of prior knowledge. Unknown parameters in this specification are estimated by probabilistic inference. Often, however, the problem size poses an important challenge and in order to render the approach feasible, specialized inference methods need to be tailored to improve the computational speed and efficiency. A. Taylan Cemgil |
ACM Multimedia | 1 |
| 2006 | A generative model for music transcriptionabstractIn this paper, we present a graphical model for polyphonic music transcription. Our model, formulated as a dynamical Bayesian network, embodies a transparent and computationally tractable approach to this acoustic analysis problem. An advantage of our approach is that it places emphasis on explicitly modeling the sound generation procedure. It provides a clear framework in which both high level (cognitive) prior information on music structure can be coupled with low level (acoustic physical) information in a principled manner to perform the analysis. The model is a special case of the, generally intractable, switching Kalman filter model. Where possible, we derive, exact polynomial time inference procedures, and otherwise efficient approximations. We argue that our generative model based approach is computationally feasible for many music applications and is readily extensible to more general auditory scene analysis scenarios. A. Taylan Cemgil, Hilbert J. Kappen, David Barber |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | A Hybrid Graphical Model for Robust Feature Extraction from VideoabstractWe consider a visual scene analysis scenario where objects (e.g. people, cars) pass through the viewing field of a static camera and need to be detected and segmented from the background. For this purpose, we introduce a hybrid dynamic Bayesian network and derive an expectation propagation (EP) algorithm for robust estimation of object shapes and appearance statistics. We demonstrate the viability of the approximation on an object detection task from real videos, where objects' smooth shapes are segmented from the background. The model is readily extendible to multi-object multi-camera scenarios and can be coupled in a transparent and consistent way with a hierarchical model for object identification under uncertainty. A. Taylan Cemgil, Wojciech Zajdel, Ben J. A. Kröse |
CVPR (1) | 1 |
| 2003 | Monte Carlo Methods for Tempo Tracking and Rhythm QuantizationabstractWe present a probabilistic generative model for timing deviations in expressive music performance. The structure of the proposed model is equivalent to a switching state space model. The switch variables correspond to discrete note locations as in a musical score. The continuous hidden variables denote the tempo. We formulate two well known music recognition problems, namely tempo tracking and automatic transcription (rhythm quantization) as filtering and maximum a posteriori (MAP) state estimation tasks. Exact computation of posterior features such as the MAP state is intractable in this model class, so we introduce Monte Carlo methods for integration and optimization. We compare Markov Chain Monte Carlo (MCMC) methods (such as Gibbs sampling, simulated annealing and iterative improvement) and sequential Monte Carlo methods (particle filters). Our simulation results suggest better results with sequential methods. The methods can be applied in both online and batch scenarios such as tempo tracking and transcription and are thus potentially useful in a number of music applications such as adaptive automatic accompaniment, score typesetting and music information retrieval. A. Taylan Cemgil, Hilbert J. Kappen |
J. Artif. Intell. Res. | 1 |
| 2001 | Tempo tracking and rhythm quantization by sequential Monte CarloabstractWe present a probabilistic generative model for timing deviations in expressive music. performance. The structure of the proposed model is equivalent to a switching state space model. We formu(cid:173) late two well known music recognition problems, namely tempo tracking and automatic transcription (rhythm quantization) as fil(cid:173) tering and maximum a posteriori (MAP) state estimation tasks. The inferences are carried out using sequential Monte Carlo in(cid:173) tegration (particle filtering) techniques. For this purpose, we have derived a novel Viterbi algorithm for Rao-Blackwellized particle fil(cid:173) ters, where a subset of the hidden variables is integrated out. The resulting model is suitable for realtime tempo tracking and tran(cid:173) scription and hence useful in a number of music applications such as adaptive automatic accompaniment and score typesetting. A. Taylan Cemgil, Hilbert J. Kappen |
NIPS | 1 |