VLDB 2026 Research / reviewers in the wild / expert
Shlomo Dubnov
dblp:89/4032
· DBLP profile ↗
47ranked-venue papers
12as first author
19since 2021 · last 2025
0000-0003-0222-1125ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Simultaneous Music Separation and Generation Using Multi-Track Latent Diffusion ModelsabstractDiffusion models have recently shown strong potential in both music generation and music source separation tasks. Although in early stages, a trend is emerging towards integrating these tasks into a single framework, as both involve generating musically aligned parts and can be seen as facets of the same generative process. In this work, we introduce a latent diffusion-based multi-track generation model capable of both source separation and multi-track music synthesis by learning the joint probability distribution of tracks sharing a musical context. Our model also enables arrangement generation by creating any subset of tracks given the others. We trained our model on the Slakh2100 dataset, compared it with an existing simultaneous generation and separation model, and observed significant improvements across objective metrics for source separation, music, and arrangement generation tasks. Sound examples are available at https://msg-ld.github.io/. Tornike Karchkhadze, Mohammad Rasool Izadi, Shlomo Dubnov |
ICASSP | 3 |
| 2025 | kNN-SVC: Robust Zero-Shot Singing Voice Conversion with Additive Synthesis and Concatenation Smoothness OptimizationabstractRobustness is critical in zero-shot singing voice conversion (SVC). This paper introduces two novel methods to strengthen the robustness of the kNN-VC framework for SVC. First, kNN-VC’s core representation, WavLM, lacks harmonic emphasis, resulting in dull sounds and ringing artifacts. To address this, we leverage the bijection between WavLM, pitch contours, and spectrograms to perform additive synthesis, integrating the resulting waveform into the model to mitigate these issues. Second, kNN-VC overlooks concatenative smoothness, a key perceptual factor in SVC. To enhance smoothness, we propose a new distance metric that filters out unsuitable kNN candidates and optimize the summing weights of the candidates during inference. Although our techniques are built on the kNN-VC framework for implementation convenience, they are broadly applicable to general concatenative neural synthesis models. Experimental results validate the effectiveness of these modifications in achieving robust SVC. Demo: http://knnsvc.com. Keren Shao, Ke Chen 0021, Matthew Baas, Shlomo Dubnov |
ICASSP | 4 |
| 2025 | Synthesizing Composite Hierarchical Structure from Symbolic Music CorporaabstractWestern music is an innately hierarchical system of interacting levels of structure, from fine-grained melody to high-level form. In order to analyze music compositions holistically and at multiple granularities, we propose a unified, hierarchical meta-representation of musical structure called the structural temporal graph (STG). For a single piece, the STG is a data structure that defines a hierarchy of progressively finer structural musical features and the temporal relationships between them. We use the STG to enable a novel approach for deriving a representative structural summary of a music corpus, which we formalize as a dually NP-hard combinatorial optimization problem. Our approach first applies simulated annealing to develop a measure of structural distance between two music pieces rooted in graph isomorphism. Our approach then combines the formal guarantees of SMT solvers with nested simulated annealing over structural distances to produce a structurally sound, representative centroid STG for an entire corpus of STGs from individual pieces. To evaluate our approach, we conduct experiments verifying that structural distance accurately differentiates between music pieces, and that derived centroids accurately structurally characterize their corpora. Ilana Shapiro, Ruanqianqian (Lisa) Huang, Zachary Novack, Cheng-i Wang, Hao-Wen Dong, Taylor Berg-Kirkpatrick, Shlomo Dubnov, Sorin Lerner |
IJCAI | 7 |
| 2024 | Interpreting Graphic Notation with MusicLDM: An AI Improvisation of Cornelius Cardew's TreatiseabstractThis work presents a novel method for composing and improvising music inspired by Cornelius Cardew’s Treatise, using AI to bridge graphic notation and musical expression. By leveraging OpenAI’s ChatGPT to interpret the abstract visual elements of Treatise, we convert these graphical images into descriptive textual prompts. These prompts are then input into MusicLDM, a pre-trained latent diffusion model designed for music generation. We introduce a technique called "outpainting," which overlaps sections of AI-generated music to create a seamless and cohesive composition. We demostrate a new perspective on performing and interpreting graphic scores, showing how AI can transform visual stimuli into sound and expand the creative possibilities in contemporary/experimental music composition. Musical pieces are available at https://bit.ly/TreatiseAI. Tornike Karchkhadze, Keren Shao, Shlomo Dubnov |
IEEE Big Data | 3 |
| 2024 | MusicLDM: Enhancing Novelty in text-to-music Generation Using Beat-Synchronous mixup StrategiesabstractDiffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited availability of music data and sensitive issues related to copyright and plagiarism. In this paper, to tackle these challenges, we first construct a state-of-the-art text-to-music model, MusicLDM, that adapts Stable Diffusion and AudioLDM architectures to the music domain. Then, to address the limitations of training data and to avoid plagiarism, we leverage a beat tracking model and propose two different mixup strategies for data augmentation: beat-synchronous audio mixup and beat-synchronous latent mixup, which recombine training audio directly or via a latent embeddings space, respectively. Such mixup strategies encourage the model to interpolate between musical training samples and generate new music within the convex hull of the training data, making the generated music more diverse while still staying faithful to the corresponding style. In addition to popular evaluation metrics, we design several new evaluation metrics based on CLAP score to demonstrate that our proposed MusicLDM and beat-synchronous mixup strategies improve both the quality and novelty of generated music, as well as the correspondence between input text and generated music. Ke Chen 0021, Yusong Wu, Haohe Liu, Marianna Nezhurina, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 6 |
| 2024 | Binaural Sound Source Localization Using a Hybrid Time and Frequency Domain ModelabstractThis paper introduces a new approach to sound source localization using head-related transfer function (HRTF) characteristics, which enable precise full-sphere localization from raw data. While previous research focused primarily on using extensive microphone arrays in the frontal plane, this arrangement often encountered limitations in accuracy and robustness when dealing with smaller microphone arrays. Our model proposes using both time and frequency domain for sound source localization while utilizing Deep Learning (DL) approach. The performance of our proposed model, surpasses the current state-of-the-art results. Specifically, it boasts an average angular error of 0.24° and an average Euclidean distance of 0.01 meters, while the known stateof-the-art gives average angular error of 19.07° and average Euclidean distance of 1.08 meters. This level of accuracy is of paramount importance for a wide range of applications, including robotics, virtual reality, and aiding individuals with cochlear implants (CI). Gil Geva, Olivier Warusfel, Shlomo Dubnov, Tammuz Dubnov, Amir Amedi, Yacov Hel-Or |
ICASSP | 3 |
| 2024 | SelfVC: Voice Conversion With Iterative Refinement using Self TransformationsabstractWe propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that separately encode speaker characteristics and linguistic content. However, disentangling speech representations to capture such attributes using task-specific loss terms can lead to information loss. In this work, instead of explicitly disentangling attributes with loss terms, we present a framework to train a controllable voice conversion model on entangled speech representations derived from self-supervised learning (SSL) and speaker verification models. First, we develop techniques to derive prosodic information from the audio signal and SSL representations to train predictive submodules in the synthesis model. Next, we propose a training strategy to iteratively improve the synthesis model for voice conversion, by creating a challenging training objective using self-synthesized examples. We demonstrate that incorporating such self-synthesized examples during training improves the speaker similarity of generated speech as compared to a baseline voice conversion model trained solely on heuristically perturbed inputs. Our framework is trained without any text and achieves state-of-the-art results in zero-shot voice conversion on metrics evaluating naturalness, speaker similarity, and intelligibility of synthesized audio. Paarth Neekhara, Shehzeen Hussain, Rafael Valle, Boris Ginsburg, Rishabh Ranjan, Shlomo Dubnov, Farinaz Koushanfar, Julian J. McAuley |
ICML | 6 |
| 2024 | Retrieval Guided Music Captioning via Multimodal Prefixes
Nikita Srivatsan, Ke Chen 0021, Shlomo Dubnov, Taylor Berg-Kirkpatrick |
IJCAI | 3 |
| 2024 | Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation
Ke Chen 0021, Jiaqi Su, Taylor Berg-Kirkpatrick, Shlomo Dubnov, Zeyu Jin |
INTERSPEECH | 4 |
| 2023 | Equipping Pretrained Unconditional Music Transformers with Instrument and Genre ControlsabstractThe “pretraining-and-finetuning” paradigm has become a norm for training domain-specific models in natural language processing and computer vision. In this work, we aim to examine this paradigm for symbolic music generation through leveraging the largest ever symbolic music dataset sourced from the MuseScore forum. We first pretrain a large unconditional transformer model using 1.5 million songs. We then propose a simple technique to equip this pretrained unconditional music transformer model with instrument and genre controls by finetuning the model with additional control tokens. Our proposed representation offers improved high-level controllability and expressiveness against two existing representations. The experimental results show that the proposed model can successfully generate music with user-specified instruments and genre. In a subjective listening test, the proposed model outperforms the pretrained baseline model in terms of coherence, harmony, arrangement and overall quality. Weihan Xu, Julian J. McAuley, Shlomo Dubnov, Hao-Wen Dong |
IEEE Big Data | 3 |
| 2023 | Multitrack Music TransformerabstractExisting approaches for generating multitrack music with transformer models have been limited in terms of the number of instruments, the length of the music segments and slow inference. This is partly due to the memory requirements of the lengthy input sequences necessitated by existing representations. In this work, we propose a new multitrack music representation that allows a diverse set of instruments while keeping a short sequence length. Our proposed Multitrack Music Transformer (MMT) achieves comparable performance with state-of-the-art systems, landing in between two recently proposed models in a subjective listening test, while achieving substantial speedups and memory reductions over both, making the method attractive for real time improvisation or near real time creative applications. Further, we propose a new measure for analyzing musical self-attention and show that the trained model attends more to notes that form a consonant interval with the current note and to notes that are 4N beats away from the current step. Hao-Wen Dong, Ke Chen 0021, Shlomo Dubnov, Julian J. McAuley, Taylor Berg-Kirkpatrick |
ICASSP | 3 |
| 2023 | Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption AugmentationabstractContrastive learning has shown remarkable success in the field of multimodal representation learning. In this paper, we propose a pipeline of contrastive language-audio pretraining to develop an audio representation by combining audio data with natural language descriptions. To accomplish this target, we first release LAION-Audio-630K, a large collection of 633,526 audio-text pairs from different data sources. Second, we construct a contrastive language-audio pretraining model by considering different audio encoders and text encoders. We incorporate the feature fusion mechanism and keyword-to-caption augmentation into the model design to further enable the model to process audio inputs of variable lengths and enhance the performance. Third, we perform comprehensive experiments to evaluate our model across three tasks: text-to-audio retrieval, zero-shot audio classification, and supervised audio classification. The results demonstrate that our model achieves superior performance in text-to-audio retrieval task. In audio classification tasks, the model achieves state-of-the-art performance in the zero-shot setting and is able to obtain performance comparable to models’ results in the non-zero-shot setting. LAION-Audio-630K1and the proposed model2are both available to the public. Yusong Wu, Ke Chen 0021, Yuchen Hui, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 6 |
| 2022 | Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled DataabstractDeep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a single model to target multiple sources, they have difficulty generalizing to unseen sources. In this paper, we propose a three-component pipeline to train a universal audio source separator from a large, but weakly-labeled dataset: AudioSet. First, we propose a transformer-based sound event detection system for processing weakly-labeled training data. Second, we devise a query-based audio separation model that leverages this data for model training. Third, we design a latent embedding processor to encode queries that specify audio targets for separation, allowing for zero-shot generalization. Our approach uses a single model for source separation of multiple sound types, and relies solely on weakly-labeled data for training. In addition, the proposed audio separator can be used in a zero-shot setting, learning to separate types of audio sources that were never seen in training. To evaluate the separation performance, we test our model on MUSDB18, while training on the disjoint AudioSet. We further verify the zero-shot performance by conducting another experiment on audio source types that are held-out from training. The model achieves comparable Source-to-Distortion Ratio (SDR) performance to current supervised models in both cases. Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
AAAI | 6 |
| 2022 | HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and DetectionabstractAudio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted in this field. However, existing audio transformers require large GPU memories and long training time, meanwhile relying on pretrained vision models to achieve high performance, which limits the model’s scalability in audio tasks. To combat these problems, we introduce HTS-AT: an audio transformer with a hierarchical structure to reduce the model size and training time. It is further combined with a token-semantic module to map final outputs into class featuremaps, thus enabling the model for the audio event detection (i.e. localization in time). We evaluate HTS-AT on three datasets of audio classification where it achieves new state-of-the-art (SOTA) results on AudioSet and ESC50, and equals the SOTA on Speech Command V2. It also achieves better performance in event localization than the previous CNN-based models. Moreover, HTS-AT requires only 35% model parameters and 15% training time of the previous audio transformer. These results demonstrate the high performance and high efficiency of HTS-AT. Ke Chen 0021, Xingjian Du, Bilei Zhu, Zejun Ma 0001, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 6 |
| 2022 | Tonet: Tone-Octave Network for Singing Melody Extraction from Polyphonic MusicabstractSinging melody extraction is an important problem in the field of music information retrieval. Existing methods typically rely on frequency-domain representations to estimate the sung frequencies. However, this design does not lead to human-level performance in the perception of melody information for both tone (pitch-class) and octave. In this paper, we propose TONet1, a plug-and-play model that improves both tone and octave perceptions by leveraging a novel input representation and a novel network architecture. First, we present an improved input representation, the Tone-CFP, that explicitly groups harmonics via a rearrangement of frequency-bins. Second, we introduce an encoder-decoder architecture that is designed to obtain a salience feature map, a tone feature map, and an octave feature map. Third, we propose a tone-octave fusion mechanism to improve the final salience feature map. Experiments are done to verify the capability of TONet with various baseline backbone models. Our results show that tone-octave fusion with Tone-CFP can significantly improve the singing voice extraction performance across various datasets – with substantial gains in octave and tone accuracy. Ke Chen 0021, Shuai Yu 0002, Cheng-i Wang, Wei Li 0012, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 6 |
| 2022 | Cross-modal Adversarial ReprogrammingabstractWith the abundance of large-scale deep learning models, it has become possible to repurpose pre-trained networks for new tasks. Recent works on adversarial reprogramming have shown that it is possible to repurpose neural networks for alternate tasks without modifying the network architecture or parameters. However these works only consider original and target tasks within the same data domain. In this work, we broaden the scope of adversarial reprogramming beyond the data modality of the original task. We analyze the feasibility of adversarially repurposing image classification neural networks for Natural Language Processing (NLP) and other sequence classification tasks. We design an efficient adversarial program that maps a sequence of discrete tokens into an image which can be classified to the desired class by an image classification model. We demonstrate that by using highly efficient adversarial programs, we can reprogram image classifiers to achieve competitive performance on a variety of text and sequence classification benchmarks without retraining the network. Paarth Neekhara, Shehzeen Hussain, Jinglong Du, Shlomo Dubnov, Farinaz Koushanfar, Julian J. McAuley |
WACV | 4 |
| 2021 | Expressive Neural Voice CloningabstractVoice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, these approaches lack the ability to control the expressiveness of synthesized audio. In this work, we propose a controllable voice cloning method that allows fine-grained control over various style aspects of the synthesized speech for an unseen speaker. We achieve this by explicitly conditioning the speech synthesis model on a speaker encoding, pitch contour and latent style tokens during training. Through both quantitative and qualitative evaluations, we show that our framework can be used for various expressive voice cloning tasks using only a few transcribed or untranscribed speech samples for a new speaker. These cloning tasks include style transfer from a reference speech, synthesizing speech directly from text, and fine-grained style control by manipulating the style conditioning variables during inference. Paarth Neekhara, Shehzeen Hussain, Shlomo Dubnov, Farinaz Koushanfar, Julian J. McAuley |
ACML | 3 |
| 2021 | Restoring Eye Contact to the Virtual Classroom with Machine LearningabstractNonverbal communication, in particular eye contact, is a critical element of the music classroom, shown to keep students on task, coordinate musical flow, and communicate improvisational ideas. Unfortunately, this nonverbal aspect to performance and pedagogy is lost in the virtual classroom. In this paper, we propose a machine learning system which uses single instance, single camera image frames as input to estimate the gaze target of a user seated in front of their computer, augmenting the user's video feed with a display of the estimated gaze target and thereby restoring nonverbal communication of directed gaze. The proposed estimation system consists of modular machine learning blocks, leading to a target-oriented (rather than coordinate-oriented) gaze prediction. We instantiate one such example of the complete system to run a pilot study in a virtual music classroom over Zoom software. Inference time and accuracy meet benchmarks for videoconferencing applications, and quantitative and qualitative results of pilot experiments include improved success of cue interpretation and student-reported formation of collaborative, communicative relationships between conductor and musician. Ross Greer, Shlomo Dubnov |
CSEDU (1) | 2 |
| 2021 | WaveGuard: Understanding and Mitigating Audio Adversarial Examples
Shehzeen Hussain, Paarth Neekhara, Shlomo Dubnov, Julian J. McAuley, Farinaz Koushanfar |
USENIX Security Symposium | 3 |
| 2019 | In Fleeting Visions: Deep Neural Music Fickle PlayabstractWe describe a series of short music pieces that is generated in a semi-improvised manner by a computer, using deep representations and temporal memory model constructed from learning a corpus of piano works by Sergei Prokofiev. Inspired by Prokofiev's piece of a similar name, this work explores imaginative capabilities of generative music machine through a series of passing-by sonic visions triggered by fleeting activations of the underlying musical network from a musician performer input. Unlike most other common machine learning and neural music compositions that explore stylistic imitation, the impetus here is to provide a rainbow of unimagined possibilities enabling creative human intervention and interaction with a complex system, realized in a series of improvisations, each with a different form, texture and character. Shlomo Dubnov |
Creativity & Cognition | 1 |
| 2019 | Adversarial Reprogramming of Text Classification Neural NetworksabstractPaarth Neekhara, Shehzeen Hussain, Shlomo Dubnov, Farinaz Koushanfar. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Paarth Neekhara, Shehzeen Hussain, Shlomo Dubnov, Farinaz Koushanfar |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Expediting TTS Synthesis with Adversarial VocodingabstractRecent approaches in text-to-speech (TTS) synthesis employ neural network strategies to vocode perceptually-informed spectrogram representations directly into listenable waveforms. Such vocoding procedures create a computational bottleneck in modern TTS pipelines. We propose an alternative approach which utilizes generative adversarial networks (GANs) to learn mappings from perceptually-informed spectrograms to simple magnitude spectrograms which can be heuristically vocoded. Through a user study, we show that our approach significantly outperforms naïve vocoding strategies while being hundreds of times faster than neural network vocoders used in state-of-the-art TTS systems. We also show that our method can be used to achieve state-of-the-art results in unsupervised synthesis of individual words of speech. Paarth Neekhara, Chris Donahue, Miller S. Puckette, Shlomo Dubnov, Julian J. McAuley |
INTERSPEECH | 4 |
| 2019 | Universal Adversarial Perturbations for Speech Recognition SystemsabstractIn this work, we demonstrate the existence of universal adversarial audio perturbations that cause mis-transcription of audio signals by automatic speech recognition (ASR) systems. We propose an algorithm to find a single quasi-imperceptible perturbation, which when added to any arbitrary speech signal, will most likely fool the victim speech recognition model. Our experiments demonstrate the application of our proposed technique by crafting audio-agnostic universal perturbations for the state-of-the-art ASR system -- Mozilla DeepSpeech. Additionally, we show that such perturbations generalize to a significant extent across models that are not available during training, by performing a transferability test on a WaveNet based ASR system. Paarth Neekhara, Shehzeen Hussain, Prakhar Pandey, Shlomo Dubnov, Julian J. McAuley, Farinaz Koushanfar |
INTERSPEECH | 4 |
| 2018 | Rethinking Recurrent Latent Variable Model for Music CompositionabstractWe present a model for capturing musical features and creating novel sequences of music, called the Convolutional-Variational Recurrent Neural Network. To generate sequential data, the model uses an encoder-decoder architecture with latent probabilistic connections to capture the hidden structure of music. Using the sequence-to-sequence model, our generative model can exploit samples from a prior distribution and generate a longer sequence of music. We compare the performance of our proposed model with other types of Neural Networks using the criteria of Information Rate that is implemented by Variable Markov Oracle, a method that allows statistical characterization of musical information dynamics and detection of motifs in a song. Our results suggest that the proposed model has a better statistical resemblance to the musical structure of the training data, which improves the creation of new sequences of music in the style of the originals. Eunjeong Stella Koh, Shlomo Dubnov, Dustin Wright 0002 |
MMSP | 2 |
| 2015 | Pattern discovery from audio recordings by Variable Markov Oracle: A music information dynamics approachabstractIn this paper, a framework for automatic pattern discovery within an audio recording is proposed. The concept of the proposed framework stems from music information dynamics and is realized by Variable Markov Oracle. Music information dynamics is the research area focusing on information theoretic measures describing musical structure and is thus closely related to the field of music pattern discovery. Variable Markov Oracle is a data structure that provides both fast retrieval of repeated sub-clips from a signal and efficient calculation of music information dynamics measures. Evaluation of the proposed framework is performed on the JKU Patterns Development Dataset with significantly improved performance of the current state of the art. Cheng-i Wang, Shlomo Dubnov |
ICASSP | 2 |
| 2014 | Variable Markov Oracle: A Novel Sequential Data Points Clustering Algorithm with Application to 3D Gesture Query-MatchingabstractIn this paper a new method, Variable Markov Oracle, for clustering time series data points is proposed. Variable Markov Oracle is based on previous results of Audio Oracle, a method of fast indexing repeating sub-clips in an audio stream. The proposed method is capable of discovering natural clusters with temporal relations without specifying the number of clusters. The discovery of inherent clusters in time series data points allows the devising of an efficient algorithm for time series query-matching. The ability of discovering clusters is demonstrated with a synthetic audio example, and an application of querying 3D skeletal gesture using the query-matching algorithm based on the proposed method is experimented with comparable result to state of the art. Cheng-i Wang, Shlomo Dubnov |
ISM | 2 |
| 2014 | Nonlinear dynamics of human creativityabstractThe nonlinear dynamical approach to creativity considers it as the processes responsible for the producing effective novelty, as well as the control mechanisms that regulate novelty production. Merely novel information display surprise and incongruity, to be sure, but they must also be meaningful. We present the Dynamical Origin of Creativity (DOC) hypothesis in application to music improvisation/ composition and poetry. The corresponding dynamical models are based on the general cognitive principles: transitivity, existence of metastable states, robustness and sensitivity to available information. Mikhail I. Rabinovich, Irma Tristan, Shlomo Dubnov |
SMC | 3 |
| 2011 | On the Information Geometry of Audio Streams With Applications to Similarity ComputingabstractThis paper proposes methods for information processing of audio streams using methods of information geometry. We lay the theoretical groundwork for a framework allowing the treatment of signal information as information entities, suitable for similarity and symbolic computing on audio signals. The theoretical basis of this paper is based on the information geometry of statistical structures representing audio spectrum features, and specifically through the bijection between the generic families of Bregman divergences and that of exponential distributions. The proposed framework, called Music Information Geometry, allows online segmentation of audio streams to metric balls where each ball represents a quasi-stationary continuous chunk of audio, and discusses methods to qualify and quantify information between entities for similarity computing. We define an information geometry that approximates a similarity metric space, redefine general notions in music information retrieval such as similarity between entities, and address methods for dealing with nonstationarity of audio signals. We demonstrate the framework on two sample applications for online audio structure discovery and audio matching. Arshia Cont, Shlomo Dubnov, Gérard Assayag |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | JoinSee: A Real-Time and Collaborative Hyper-Media System for Participatory Performances in the Opera of MeaningabstractIn this paper, we propose a real-time and collaborative hyper-media system that introduces database-enhanced collaborative models and multimedia processing models for creating improvised performances in the Opera of Meaning. The system provides a main story media and a corresponding shared canvas that is shared among Internet-wide user communities. Our shared canvas mechanisms make it possible to describe and share users' ideas and impressions about the main story media. The key technology of this system is a timeline-dependent and script-driven live performance engine, which provides users with ECA rules to express and characterize the users' ideas and impressions by using existing multimedia data such as video files and image files. The system provides directors and participants of improvised performance with a set of database operators for controlling and contributing to the performance. The system motivates users to contribute to the performance by exploiting users' own media libraries and existing web services. We have implemented the prototype system which is applicable to the existing video and image files on the Web. Shuichi Kurabayashi, Shlomo Dubnov, Yasushi Kiyoki |
EJC | 2 |
| 2009 | Analyzing several musical instrument tones using the randomly modulated periodicity model
Shlomo Dubnov, Melvin J. Hinich |
Signal Process. | 1 |
| 2008 | Opera of Meaning: film and music performance with semantic associative searchabstractRecently artists are exploring ways for incorporating large amounts of information and networking as part of their medium. One of the main challenges in applying information technology to film and opera is in relating different types of media to the meaning of story narrative. Opera of Meaning is a new format for distributed, collaborative and interactive viewing where the association of different media elements is done dynamically by semantic and impression search that is performed by the public during the performance in context of a main story. This opens new research questions in database modeling and semantic technology related to story meaning, media auto-tagging, automatic editing and mixing, user interaction, social networking and more. We plan to offer this format to artists, producers and the public, opening a new venue for social creation and experiencing of impression and meaning in digital media. Shlomo Dubnov, Yasushi Kiyoki |
EJC | 1 |
| 2008 | Unified View of Prediction and Repetition Structure in Audio Signals With Application to Interest Point DetectionabstractIn this paper, we present a new method for analysis of musical structure that captures local prediction and global repetition properties of audio signals in one information processing framework. The method is motivated by a recent work in music perception where machine features were shown to correspond to human judgments of familiarity and emotional force when listening to music. Using a notion of information rate in a model-based framework, we develop a measure of mutual information between past and present in a time signal and show that it consist of two factors - prediction property related to data statistics within an individual block of signal features, and repetition property based on differences in model likelihood across blocks. The first factor, when applied to spectral representation of audio signals, is known as spectral anticipation, and the second factor is known as recurrence analysis. We present algorithms for estimation of these measures and create a visualization that displays their temporal structure in musical recordings. Considering these features as a measure of the amount of information processing that a listening system performs on a signal, information rate is used to detect interest points in music. Several musical works with different performances are analyzed in this paper, and their structure and interest points are displayed and discussed. Extensions of this approach towards a general framework of characterizing machine listening experience are suggested. Shlomo Dubnov |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Computer audition: an introduction and research surveyabstractNo abstract available. Shlomo Dubnov |
ACM Multimedia | 1 |
| 2006 | Structural and affective aspects of music from statistical audio signal analysisabstractAbstract Understanding and modeling human experience and emotional response when listening to music are important for better understanding of the stylistic choices in musical composition. In this work, we explore the relation of audio signal structure to human perceptual and emotional reactions. Memory, repetition, and anticipatory structure have been suggested as some of the major factors in music that might influence and possibly shape these responses. The audio analysis was conducted on two recordings of an extended contemporary musical composition by one of the authors. Signal properties were analyzed using statistical analyses of signal similarities over time and information theoretic measures of signal redundancy. They were then compared to Familiarity Rating and Emotional Force profiles, as recorded continually by listeners hearing the two versions of the piece in a live‐concert setting. The analysis shows strong evidence that signal properties and human reactions are related, suggesting applications of these techniques to music understanding and music information‐retrieval systems. Shlomo Dubnov, Stephen McAdams, Roger Reynolds |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2006 | Generalized likelihood ratio test for voiced-unvoiced decision in noisy speech using the harmonic modelabstractIn this paper, a novel method for voiced-unvoiced decision within a pitch tracking algorithm is presented. Voiced-unvoiced decision is required for many applications, including modeling for analysis/synthesis, detection of model changes for segmentation purposes and signal characterization for indexing and recognition applications. The proposed method is based on the generalized likelihood ratio test (GLRT) and assumes colored Gaussian noise with unknown covariance. Under voiced hypothesis, a harmonic plus noise model is assumed. The derived method is combined with a maximum a-posteriori probability (MAP) scheme to obtain a pitch and voicing tracking algorithm. The performance of the proposed method is tested using several speech databases for different levels of additive noise and phone speech conditions. Results show that the GLRT is robust to speaker and environmental conditions and performs better than existing algorithms. Etan Fisher, Joseph Tabrikian, Shlomo Dubnov |
IEEE Trans. Speech Audio Process. | 3 |
| 2004 | A method for directionally-disjoint source separation in convolutive environmentabstractWe propose a new method for source separation that is based on directionally- disjoint estimation of the transfer functions between microphones and sources at different frequencies and at multiple times. The directions are estimated from eigenvectors of the microphones' correlation matrix. Smoothing and association of transfer function parameters across different frequencies is achieved by simultaneous Kalman filtering of the noisy amplitude and phase estimates. This approach allows estimating transfer functions even in the case where the difference between the sources is in delay only and it can operate both for wideband and narrowband sources. Simulation results show superior performance in comparison to other, existing methods. Shlomo Dubnov, Joseph Tabrikian, Miki Arnon-Targan |
ICASSP (5) | 1 |
| 2004 | Using Factor Oracles for Machine Improvisation
Gérard Assayag, Shlomo Dubnov |
Soft Comput. | 2 |
| 2004 | Generalization of spectral flatness measure for non-Gaussian linear processesabstractWe present an information-theoretic measure for the amount of randomness or stochasticity that exists in a signal. This measure is formulated in terms of the rate of growth of multi-information for every new signal sample of the signal that is observed over time. In case of a Gaussian statistics it is shown that this measure is equivalent to the well-known spectral flatness measure that is commonly used in audio processing. For nonGaussian linear processes a generalized spectral flatness measure is developed, which estimates the excessive structure that is present in the signal due to the nonGaussianity of the innovation process. An estimator for this measure is developed using Negentropy approximation to the non-Gaussian signal and the innovation process statistics. Applications of this new measure are demonstrated for the problem of voiced/unvoiced determination, showing improved performance. Shlomo Dubnov |
IEEE Signal Process. Lett. | 1 |
| 2004 | Maximum a-posteriori probability pitch tracking in noisy environments using harmonic modelabstractModern speech processing applications require operation on signal of interest that is contaminated by high level of noise. This situation calls for a greater robustness in estimation of the speech parameters, a task which is hard to achieve using standard speech models. In this paper, we present an optimal estimation procedure for sound signals (such as speech) that are modeled by harmonic sources. The harmonic model achieves more robust and accurate estimation of voiced speech parameters. Using maximum a posteriori probability framework, successful tracking of pitch parameters is possible in ultra low signal to noise conditions (as low as -15 dB). The performance of the method is evaluated using the Keele pitch detection database with realistic background noise. The results show best performance in comparison to other state-of-the-art pitch detectors. Application of the proposed algorithm in a simple speaker identification system shows significant improvement in the performance. Joseph Tabrikian, Shlomo Dubnov, Yulya Dickalov |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Improved low bit-rate audio compression using reduced rank ICA instead of psychoacoustic modelingabstractTraditional audio coding is based on a perceptual compression paradigm that exploits psychoacoustic information to efficiently encode audio signals. Recently, extensive research has been conducted in order to understand how the brain encodes natural signals. These results suggest that the encoding process is very efficient in terms of redundancy reduction of the signal information. It could be that the psychoacoustic effects (such as the masking effect) are only a special case of a more general redundancy reduction mechanism that exists in the auditory pathway. Motivated by this work we propose a new audio coding scheme that is based on improved sound representation found by independent component analysis. Using a local linear, low rank, nonorthogonal transform, we remove additional redundancies in the signal. At low bitrates this coding scheme gives results superior to a legacy perceptual encoding scheme for different kinds of audio signals. Adiel Ben-Shalom, Michael Werman, Shlomo Dubnov |
ICASSP (5) | 3 |
| 2003 | Generalized likelihood ratio test for voiced/unvoiced decision using the harmonic plus noise modelabstractIn this paper, a novel method for voiced/invoiced decision in speech and music signals is presented. Voiced/unvoiced decision is required for many applications, including better modeling for analysis/synthesis, detection of model changes for segmentation purposes and better signal characterization for indexing and recognition applications. The proposed method is based on the generalized likelihood ratio test (GLRT) and assumes colored Gaussian noise with unknown covariance. Under voiced hypothesis, a harmonic plus noise model is assumed. The derived method is combined with a maximum a-posteriori probability (MAP) scheme to obtain a voiced unvoiced tracking algorithm. The performance of the proposed method is tested under the Keele University database for different signal-to-noise ratios (SNRs), and the results show that the algorithm performs well even under severe noise conditions. Etan Fisher, Joseph Tabrikian, Shlomo Dubnov |
ICASSP (1) | 3 |
| 2002 | Speech enhancement by harmonic modeling via map pitch trackingabstractIn this paper we present a procedure for estimating the parameters of speech signals that are contaminated by high level of noise. The proposed estimation method is developed by assuming a harmonic model for the voiced frame hypothesis. A Maximum A-posteriori Probability tracking method is developed for estimating time-varying pitch. Signal recoIl8truction is achieved by projecting the signal onto the subspace of harmonic signals with the optimal estimates of the fundamental frequency. The performance of the proposed method is evaluated and compared to other existing methods using a large pitch detection database. It is shown that the proposed method for pitch estimation is more robust and much more accurate in terms of mean-square-error and gross error rate, in comparison to other existing methods, specially at ultra low signal-to-noise ratios (as low as −15 dB). Examples of speech reconstruction/enhancement are also presented in the paper. Joseph Tabrikian, Shlomo Dubnov, Yulya Dickalov |
ICASSP | 2 |
| 2002 | Robust temporal and spectral modeling for query By melodyabstractQuery by melody is the problem of retrieving musical performances from melodies. Retrieval of real performances is complicated due to the large number of variations in performing a melody and the presence of colored accompaniment noise. We describe a simple yet effective probabilistic model for this task. We describe a generative model that is rich enough to capture the spectral and temporal variations of musical performances and allows for tractable melody retrieval. While most of previous studies on music retrieval from melodies were performed with either symbolic (e.g. MIDI) data or with monophonic (single instrument) performances, we performed experiments in retrieving live and studio recordings of operas that contain a leading vocalist and rich instrumental accompaniment. Our results show that the probabilistic approach we propose is effective and can be scaled to massive datasets. Shai Shalev-Shwartz, Shlomo Dubnov, Nir Friedman, Yoram Singer |
SIGIR | 2 |
| 2002 | A New Nonparametric Pairwise Clustering Algorithm Based on Iterative Estimation of Distance Profiles
Shlomo Dubnov, Ran El-Yaniv, Yoram Gdalyahu, Elad Schneidman, Naftali Tishby, Golan Yona |
Mach. Learn. | 1 |
| 1998 | Study of spectro-temporal parameters in musical performance for expressive instrument synthesisabstractModeling expressive synthesis requires one to deal with a wide range of sound's behaviors that include spectral dynamics which are idiomatic to the instrument or characteristics of the playing technique. In the paper we analyze short time evolution of power-spectral envelopes, adaptively changing the analysis parameters so as to variably include filter envelopes, and the pitched excitation parts. Typical sequences of sound features are obtained from quantized sequences of parameters. By applying string matching techniques, new sequences of features are created and new sounds that have dynamically changing temporal behavior are re-synthesized. Shlomo Dubnov, Xavier Rodet |
SMC | 1 |
| 1997 | Analysis of sound textures in musical and machine sounds by means of higher order statistical featuresabstractIn this paper we describe a sound classification method, which seems to be applicable to a broad domain of stationary, non-musical sounds, such as machine noises and other man made non-periodic sounds. The method is based on matching higher order spectra (HOS) of the acoustic signals and it generalizes our earlier results on classification of sustained musical sounds by higher order moments. An efficient "decorrelated matched filter" implemetation is presented. The results show good sound classification statistics and a comparison to spectral matching methods is also discussed. Shlomo Dubnov, Naftali Tishby |
ICASSP | 1 |
| 1994 | Acoustic spectral estimation using higher order statisticsabstractAssuming an autoregressive (AR) filter model driven by a non-Gaussian white noise, we formulate a general parameter estimation problem. A maximum likelihood solution gives an AR estimate of the filter and the probability distribution function parameters for non-Gaussian input. The proposed method is optimal in the information theoretic sense, giving the most probable model for the source and filter under the higher order statistics constrains of the observed signal. Analysis of human singing voices and musical instruments is presented and its acoustic interpretation is discussed. Shlomo Dubnov, Naftali Tishby |
ICPR (3) | 1 |