Shlomo Dubnov

dblp:89/4032 · DBLP profile ↗
← Back
6ranked-venue papers in the field
2as first author
2since 2021 · last 2024
0000-0003-0222-1125ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 2 (1 first)Big Data, Cloud & Distributed Data Systems · 2Knowledge Engineering, Semantic Web & Information Systems · 2 (1 first)
YearPublicationVenuePosition
2024 Interpreting Graphic Notation with MusicLDM: An AI Improvisation of Cornelius Cardew's Treatise
abstract
This work presents a novel method for composing and improvising music inspired by Cornelius Cardew’s Treatise, using AI to bridge graphic notation and musical expression. By leveraging OpenAI’s ChatGPT to interpret the abstract visual elements of Treatise, we convert these graphical images into descriptive textual prompts. These prompts are then input into MusicLDM, a pre-trained latent diffusion model designed for music generation. We introduce a technique called "outpainting," which overlaps sections of AI-generated music to create a seamless and cohesive composition. We demostrate a new perspective on performing and interpreting graphic scores, showing how AI can transform visual stimuli into sound and expand the creative possibilities in contemporary/experimental music composition. Musical pieces are available at https://bit.ly/TreatiseAI.
Tornike Karchkhadze, Keren Shao, Shlomo Dubnov
IEEE Big Data3
2023 Equipping Pretrained Unconditional Music Transformers with Instrument and Genre Controls
abstract
The “pretraining-and-finetuning” paradigm has become a norm for training domain-specific models in natural language processing and computer vision. In this work, we aim to examine this paradigm for symbolic music generation through leveraging the largest ever symbolic music dataset sourced from the MuseScore forum. We first pretrain a large unconditional transformer model using 1.5 million songs. We then propose a simple technique to equip this pretrained unconditional music transformer model with instrument and genre controls by finetuning the model with additional control tokens. Our proposed representation offers improved high-level controllability and expressiveness against two existing representations. The experimental results show that the proposed model can successfully generate music with user-specified instruments and genre. In a subjective listening test, the proposed model outperforms the pretrained baseline model in terms of coherence, harmony, arrangement and overall quality.
Weihan Xu, Julian J. McAuley, Shlomo Dubnov, Hao-Wen Dong
IEEE Big Data3
2009 JoinSee: A Real-Time and Collaborative Hyper-Media System for Participatory Performances in the Opera of Meaning
abstract
In this paper, we propose a real-time and collaborative hyper-media system that introduces database-enhanced collaborative models and multimedia processing models for creating improvised performances in the Opera of Meaning. The system provides a main story media and a corresponding shared canvas that is shared among Internet-wide user communities. Our shared canvas mechanisms make it possible to describe and share users' ideas and impressions about the main story media. The key technology of this system is a timeline-dependent and script-driven live performance engine, which provides users with ECA rules to express and characterize the users' ideas and impressions by using existing multimedia data such as video files and image files. The system provides directors and participants of improvised performance with a set of database operators for controlling and contributing to the performance. The system motivates users to contribute to the performance by exploiting users' own media libraries and existing web services. We have implemented the prototype system which is applicable to the existing video and image files on the Web.
Shuichi Kurabayashi, Shlomo Dubnov, Yasushi Kiyoki
EJC2
2008 Opera of Meaning: film and music performance with semantic associative search
abstract
Recently artists are exploring ways for incorporating large amounts of information and networking as part of their medium. One of the main challenges in applying information technology to film and opera is in relating different types of media to the meaning of story narrative. Opera of Meaning is a new format for distributed, collaborative and interactive viewing where the association of different media elements is done dynamically by semantic and impression search that is performed by the public during the performance in context of a main story. This opens new research questions in database modeling and semantic technology related to story meaning, media auto-tagging, automatic editing and mixing, user interaction, social networking and more. We plan to offer this format to artists, producers and the public, opening a new venue for social creation and experiencing of impression and meaning in digital media.
Shlomo Dubnov, Yasushi Kiyoki
EJC1
2006 Structural and affective aspects of music from statistical audio signal analysis
abstract
Abstract Understanding and modeling human experience and emotional response when listening to music are important for better understanding of the stylistic choices in musical composition. In this work, we explore the relation of audio signal structure to human perceptual and emotional reactions. Memory, repetition, and anticipatory structure have been suggested as some of the major factors in music that might influence and possibly shape these responses. The audio analysis was conducted on two recordings of an extended contemporary musical composition by one of the authors. Signal properties were analyzed using statistical analyses of signal similarities over time and information theoretic measures of signal redundancy. They were then compared to Familiarity Rating and Emotional Force profiles, as recorded continually by listeners hearing the two versions of the piece in a live‐concert setting. The analysis shows strong evidence that signal properties and human reactions are related, suggesting applications of these techniques to music understanding and music information‐retrieval systems.
Shlomo Dubnov, Stephen McAdams, Roger Reynolds
J. Assoc. Inf. Sci. Technol.1
2002 Robust temporal and spectral modeling for query By melody
abstract
Query by melody is the problem of retrieving musical performances from melodies. Retrieval of real performances is complicated due to the large number of variations in performing a melody and the presence of colored accompaniment noise. We describe a simple yet effective probabilistic model for this task. We describe a generative model that is rich enough to capture the spectral and temporal variations of musical performances and allows for tractable melody retrieval. While most of previous studies on music retrieval from melodies were performed with either symbolic (e.g. MIDI) data or with monophonic (single instrument) performances, we performed experiments in retrieving live and studio recordings of operas that contain a leading vocalist and rich instrumental accompaniment. Our results show that the probabilistic approach we propose is effective and can be scaled to massive datasets.
Shai Shalev-Shwartz, Shlomo Dubnov, Nir Friedman, Yoram Singer
SIGIR2