VLDB 2026 Research / reviewers in the wild / expert
Roger B. Dannenberg
dblp:59/6414
· DBLP profile ↗
29ranked-venue papers
15as first author
8since 2021 · last 2025
0000-0003-1823-9856ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-authorComputer networks · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech ReferenceabstractWe propose a unified framework for Singing Voice Synthesis (SVS) and Conversion (SVC), addressing the limitations of existing approaches in cross-domain SVS/SVC, poor output musicality, and scarcity of singing data. Our framework enables control over multiple aspects, including language content based on lyrics, performance attributes based on a musical score, singing style and vocal techniques based on a selector, and voice identity based on a speech sample. The proposed zero-shot learning paradigm consists of one SVS model and two SVC models, utilizing pre-trained content embeddings and a diffusion-based generator. The proposed framework is also trained on mixed datasets comprising both singing and speech audio, allowing singing voice cloning based on speech reference. Experiments show substantial improvements in timbre similarity and musicality over state-of-the-art baselines, providing insights into other low-data music tasks such as instrumental style transfer. Examples can be found at: everyone-can-sing.github.io. Shuqi Dai, Roger B. Dannenberg, Zeyu Jin |
ICASSP | 3 |
| 2025 | Algorithmic Composition Using Narrative Structure and TensionabstractThis paper describes an approach to algorithmic music composition that takes narrative structures as input, allowing composers to create music directly from narrative elements. Creating narrative development in music remains a challenging task in algorithmic composition. Our system addresses this by combining leitmotifs to represent characters, generative grammars for harmonic coherence, and evolutionary algorithms to align musical tension with narrative progression. The system operates at different scales, from overall plot structure to individual motifs, enabling both autonomous composition and co-creation with varying degrees of user control. Evaluation with compositions based on tales demonstrated the system's ability to compose music that supports narrative listening and aligns with its source narratives, while being perceived as familiar and enjoyable. Francisco Braga, Gilberto Bernardes, Roger B. Dannenberg, Nuno Correia 0001 |
IJCAI | 3 |
| 2025 | Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent RecognitionabstractSinging accent research is underexplored compared to speech accent studies, primarily due to the scarcity of suitable datasets. Existing singing datasets often suffer from detail loss, frequently resulting from the vocal-instrumental separation process. Additionally, they often lack regional accent annotations. To address this, we introduce the Multi-Accent Mandarin Dry-Vocal Singing Dataset (MADVSD). MADVSD comprises over 670 hours of dry vocal recordings from 4,206 native Mandarin speakers across nine distinct Chinese regions. In addition to each participant recording audio of three popular songs in their native accent, they also recorded phonetic exercises covering all Mandarin vowels and a full octave range. We validated MADVSD through benchmark experiments in singing accent recognition, demonstrating its utility for evaluating state-of-the-art speech models in singing contexts. Furthermore, we explored dialectal influences on singing accent and analyzed the role of vowels in accentual variations, leveraging MADVSD's unique phonetic exercises. Shulei Ji, Le Ma 0002, Yuhang Jin, Shun Lei, Jianyi Chen, Haoying Fu, Roger B. Dannenberg |
ACM Multimedia | 8 |
| 2025 | Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
Ruibin Yuan, Ziqi Geng, Hengjia Li, Xingwei Qu, Songye Chen, Haoying Fu, Roger B. Dannenberg |
ACM Multimedia | 9 |
| 2024 | MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised TrainingabstractSelf-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Although SSL has been proven effective in speech and audio, its application to music audio has yet to be thoroughly explored. This is partially due to the distinctive challenges associated with modelling musical knowledge, particularly tonal and pitched characteristics of music.
To address this research gap, we propose an acoustic **M**usic und**ER**standing model with large-scale self-supervised **T**raining (**MERT**), which incorporates teacher models to provide pseudo labels in the masked language modelling (MLM) style acoustic pre-training.
In our exploration, we identified an effective combination of teacher models, which outperforms conventional speech and audio approaches in terms of performance.
This combination includes an acoustic teacher based on Residual Vector Quantization - Variational AutoEncoder (RVQ-VAE) and a musical teacher based on the Constant-Q Transform (CQT).
Furthermore, we explore a wide range of settings to overcome the instability in acoustic language model pre-training, which allows our designed paradigm to scale from 95M to 330M parameters.
Experimental results indicate that our model can generalise and perform well on 14 music understanding tasks and attain state-of-the-art (SOTA) overall scores. Ruibin Yuan, Ge Zhang 0009, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin 0002, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi 0001, Wenhao Huang 0001, Yike Guo, Jie Fu 0001 |
ICLR | 12 |
| 2024 | The Beauty of Repetition: An Algorithmic Composition Model With Motif-Level Repetition Generator and Outline-to-Music Generator in Symbolic Music GenerationabstractMost musical compositions utilize repetition as a fundamental element to create captivating aesthetic experiences. However, the potential of repetition in machine-learning-based algorithmic composition has not been thoroughly investigated. This article aims to make an initial attempt at repetition modeling by generating motif-level repetitions and integrating them into music through a combination of example-based and domain knowledge–based learning techniques. The article presents a new Motif-to-music Generation Model (MGM) that combines a motif-level repetition generator (MRG) and an outline-to-music generator (O2MG). To train this model, a new music repetition dataset (MRD) has been created, which includes 584,329 samples from various categories of motif repetition and 3,545 outline-music sequences from pop piano music. The MRG uses a Transformer encoder to learn the representation of music notes from MRD, while the repetition-aware learner in MRG takes advantage of the unique characteristics of repetitions based on music theory. The O2MG applies a novel outline-to-music learning strategy to learn the relationships among motif-level repetitions in the music and generate music based on these repetitions. The experiments show that MGM can generate a variety of beautiful repetitions with any given motif, improving the music quality and structure of machine-composed music. Zhejing Hu, Xiao Ma 0023, Yan Liu 0004, Gong Chen 0006, Yongxu Liu 0003, Roger B. Dannenberg |
IEEE Trans. Multim. | 6 |
| 2023 | Transplayer: Timbre Style Transfer with Flexible Timbre ControlabstractMusic timbre style transfer aims at replacing the instrument timbre in a solo recording with another instrument, while preserving the musical content. Existing GAN-based methods can only achieve timbre style transfer between two given timbres. Inspired by the practice in voice conversion, we propose TransPlayer, which uses an autoencoder model with one-hot representations of instruments as the condition, and a Diffwave model trained especially for music synthesis. We evaluate our model in both the one-to-one transfer task and the many-to-many transfer task. The results prove that our method is able to provide one-to-one style transfer outputs comparable with the existing GAN-based method, and can transfer among multiple timbres with only one single model. Xinlu Liu, Yi Wang 0033, Roger B. Dannenberg |
ICASSP | 5 |
| 2023 | MARBLE: Music Audio Representation Benchmark for Universal EvaluationabstractIn the era of extensive intersection between art and Artificial Intelligence (AI), such as image generation and fiction co-creation, AI for music remains relatively nascent, particularly in music understanding. This is evident in the limited work on deep music representations, the scarcity of large-scale datasets, and the absence of a universal and community-driven benchmark. To address this issue, we introduce the Music Audio Representation Benchmark for universaL Evaluation, termed MARBLE. It aims to provide a benchmark for various Music Information Retrieval (MIR) tasks by defining a comprehensive taxonomy with four hierarchy levels, including acoustic, performance, score, and high-level description. We then establish a unified protocol based on 18 tasks on 12 public-available datasets, providing a fair and standard assessment of representations of all open-sourced pre-trained models developed on music recordings as baselines. Besides, MARBLE offers an easy-to-use, extendable, and reproducible suite for the community, with a clear statement on copyright issues on datasets. Results suggest recently proposed large-scale pre-trained musical language models perform the best in most tasks, with room for further improvement. The leaderboard and toolkit repository are published to promote future music AI research. Ruibin Yuan, Yinghao Ma, Ge Zhang 0009, Xingran Chen, Hanzhi Yin, Le Zhuo, Zeyue Tian, Binyue Deng, Ningzhi Wang, Chenghua Lin 0002, Emmanouil Benetos, Anton Ragni, Norbert Gyenge, Roger B. Dannenberg, Wenhu Chen, Gus Xia, Wei Xue 0002, Shi Wang 0002, Ruibo Liu, Yike Guo, Jie Fu 0001 |
NeurIPS | 17 |
| 2020 | Artificial Creative Intelligence: Breaking the Imitation Barrier
Rowland Chen, Roger B. Dannenberg, Bhiksha Raj, Rita Singh |
ICCC | 2 |
| 2019 | Review of "The Haskell School of Music: from Signals to Symphonies, " by Paul Hudak and Donya Quick, Cambridge University Press, 2018abstractComputer languages for music have a long and interesting history.Music signal processing in particular inspired some early work on dataflow computers, and functional programming concepts are frequently associated with music because of the importance of event sequences and sample sequences in music representation.This book offers a particular approach to music computation using the Haskell programming language and two libraries written in Haskell: Euterpea for representing musical structures and Haskell School of Music for both music performance and simple graphical interfaces for control.This book is intended to teach computer music concepts along with Haskell.The many programming examples are almost exclusively drawn from music.For example, the first Haskell type expression that appears is PitchClass, the set of music note names.Similarly, functions are introduced for musical transposition, sequences are introduced through simple melodies, and so on.Some of the powerful features of Haskell are polymorphic and higher order functions, allowing music representations and operations on them to be parameterized by the details of music events.Music can be an organization of pitches, of chords, of sound qualities, etc., and abstraction can help to avoid being too rigid and limiting.On the other hand, readers are quickly forced to reckon with these abstractions.I expect beginners will struggle with recursive definitions and parametric types.For seasoned functional programmers, there is certainly some beauty in the terse yet flexible notation offered by Euterpea.One of the attractions of computers for musicians is the ability to describe music as a dynamic process or computation rather than a static score.After introducing Euterpea, this book offers examples of algorithmic composition based on recursion, combinatorics, self-similar sequences, phasing, L-systems, and Markov chains.These are interleaved with more advanced Haskell topics such as type classes, monads, and induction proofs.Most of this book is concerned with symbolic and parametric representations involving discrete notes with symbolic or numerical parameters such as pitch and loudness.Beginning with Chapter 18, this book turns to audio processing and music synthesis, which is also supported by the Euterpea library.Some basic synthesis methods are covered, including additive synthesis, subtractive synthesis, types of modulation, and simple physical models.These sound generation capabilities mean that one can write algorithms to compose music in terms of symbolic "note" events, describe the translation or "performance" of these events into detailed and nuanced control signals, and finally convert the parametric control into audio, all within one programming language framework.This is unusual, as most music programming systems focus on one representation or abstraction Roger B. Dannenberg |
J. Funct. Program. | 1 |
| 2019 | O2: A Network Protocol for Music SystemsabstractO2 is a communication protocol for music systems that extends and interoperates with the popular Open Sound Control (OSC) protocol. Many computer musicians routinely deal with problems of interconnection, unreliable message delivery, and clock synchronization. O2 solves these problems, offering named services, automatic network address discovery, clock synchronization, and a reliable message delivery option, as well as interoperability with existing OSC libraries and applications. Aside from these new features, O2 owes much of its design to OSC, making it easy to migrate existing OSC applications to O2 or for developers familiar with OSC to begin using O2. O2 addresses the problems of interprocess communication within distributed music applications. Roger B. Dannenberg |
Wirel. Commun. Mob. Comput. | 1 |
| 2011 | A Vision of Creative Computation in Music Performance
Roger B. Dannenberg |
ICCC | 1 |
| 2010 | The Intelligent Music Editor: Towards an Automated Platform for Music Analysis and Editing
Roger B. Dannenberg, Lianhong Cai |
ICIC (2) | 2 |
| 2010 | As-If Infinitely Ranged Integer ModelabstractIntegers represent a growing and underestimated source of vulnerabilities in C and C++ programs. This paper presents the As-if Infinitely Ranged (AIR) Integer model for eliminating vulnerabilities resulting from integer overflow, truncation, and unanticipated wrapping. The AIR Integer model either produces a value equivalent to that obtained using infinitely ranged integers or results in a runtime-constraint violation. With the exception of wrapping (which is optional), this model can be implemented by a C99-conforming compiler and used by the programmer with little or no change to existing source code. Fuzz testing of libraries that have been compiled using a prototype AIR integer compiler has been effective in discovering vulnerabilities in software with low false positive and false negative rates. Furthermore, the runtime overhead of the AIR Integer model is low enough that typical applications can enable it in deployed systems for additional runtime protection. Roger B. Dannenberg, Will Dormann, David Keaton, Robert C. Seacord, David Svoboda, Alex Volkovitsky, Timothy Wilson, Thomas Plum 0003 |
ISSRE | 1 |
| 2007 | David Temperley, Music and Probability , MIT Press (2007)
Roger B. Dannenberg |
Artif. Intell. | 1 |
| 2007 | A comparative evaluation of search techniques for query-by-humming using the MUSART testbedabstractAbstract Query‐by‐humming systems offer content‐based searching for melodies and require no special musical training or knowledge. Many such systems have been built, but there has not been much useful evaluation and comparison in the literature due to the lack of shared databases and queries. The MUSART project testbed allows various search algorithms to be compared using a shared framework that automatically runs experiments and summarizes results. Using this testbed, the authors compared algorithms based on string alignment, melodic contour matching, a hidden Markov model, n‐grams, and CubyHum. Retrieval performance is very sensitive to distance functions and the representation of pitch and rhythm, which raises questions about some previously published conclusions. Some algorithms are particularly sensitive to the quality of queries. Our queries, which are taken from human subjects in a realistic setting, are quite difficult, especially for n‐gram models. Finally, simulations on query‐by‐humming performance as a function of database size indicate that retrieval performance falls only slowly as the database size increases. Roger B. Dannenberg, William P. Birmingham, Bryan Pardo, Colin Meek, George Tzanetakis |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2006 | David Cope, Computer Models of Musical Creativity, MIT Press (2005)
Roger B. Dannenberg |
Artif. Intell. | 1 |
| 2006 | Bootstrap learning for accurate onset detection
Roger B. Dannenberg |
Mach. Learn. | 2 |
| 1994 | Automated Accompaniment of Musical Ensembles
Lorin Grubb, Roger B. Dannenberg |
AAAI | 2 |
| 1993 | Real-time issues in computer musicabstractReal-time systems are essential to computer music, and new developments in scheduling, operating systems, and languages are certain to find applications in music. At the same time, we believe the issues that arise in computer music should be of interest to designers of real-time systems. Computer music systems are demanding and complex, and require a large degree of flexibility to support the creative goals of musicians. Musicians and composers were working in the "real-time business" long before the computer, and we think the art and science of music-making has much to offer.> Roger B. Dannenberg, David H. Jameson |
RTSS | 1 |
| 1993 | Tactus: Toolkit-Level Support for Synchronized Interactive Multimedia
Roger B. Dannenberg, Thomas P. Neuendorffer, Joseph M. Newcomer, Dean Rubine, David B. Anderson |
Multim. Syst. | 1 |
| 1992 | Tactus: Toolkit-Level Support for Synchronized Interactive Multimedia
Roger B. Dannenberg, Thomas P. Neuendorffer, Joseph M. Newcomer, Dean Rubine |
NOSSDAV | 1 |
| 1990 | A Structure for Efficient Update, Incremental Redisplay and Undo in Graphical EditorsabstractAbstract The design of a graphical editor requires a solution to a number of problems, including how to (1) support incremental redisplay, (2) control the granularity of display updates, (3) provide efficient access and modification to the underlying data structure, (4) handle multiple views of the same data and (5) support Undo operations. It is most important that these problems be solved without sacrificing program modularity. A new data structure, called an ItemList, provides a solution to these problems. ItemLists maintain both multiple views and multiple versions of data to simplify Undo operations and to support incremental display updates. The implementation of ItemLists is described and the use of ItemLists to create graphical editors is presented. Roger B. Dannenberg |
Softw. Pract. Exp. | 1 |
| 1989 | A gesture based user interface prototyping systemabstractGID, for Gestural Interface Designer, is an experimental system for prototyping gesture-based user interfaces. GID structures an interface as a collection of “controls”: objects that maintain an image on the display and respond to input from pointing and gesture-sensing devices. GID includes an editor for arranging controls on the screen and saving screen layouts to a file. Once an interface is created, GID provides mechanisms for routing input to the appropriate destination objects even when input arrives in parallel from several devices. GID also provides low level feature extraction and gesture representation primitives to assist in parsing gestures. Roger B. Dannenberg, D. Amon |
UIST | 1 |
| 1989 | Creating graphical interactive application objects by demonstrationabstractThe Lapidary user interface tool allows all pictorial aspects of programs to be specified graphically. In addition, the behavior of these objects at run-time can be specified using dialogue boxes and by demonstration. In particular, Lapidary allows the designer to draw pictures of application-specific graphical objects which will be created and maintained at run-time by the application. This includes the graphical entities that the end user will manipulate (such as the components of the picture), the feedback that shows which objects are selected (such as small boxes on the sides and corners of an object), and the dynamic feedback objects (such as hair-line boxes to show where an object is being dragged). In addition, Lapidary supports the construction and use of “widgets” (sometimes called interaction techniques or gadgets) such as menus, scroll bars, buttons and icons. Lapidary therefore supports using a pre-defined library of widgets, and defining a new library with a unique “look and feel.” The run-time behavior of all these objects can be specified in a straightforward way using constraints and abstract descriptions of the interactive response to the input devices. Lapidary generalizes from the specific example pictures to allow the graphics and behaviors to be specified by demonstration. Brad A. Myers, Bradley T. Vander Zanden, Roger B. Dannenberg |
UIST | 3 |
| 1986 | The computer as musical accompanistabstractArticle The computer as musical accompanist Share on Authors: W. Buxton Computer Systems Research Institute, University of Toronto, Toronto, Ontario, Canada M5S 1A4 Computer Systems Research Institute, University of Toronto, Toronto, Ontario, Canada M5S 1A4View Profile , R. Dannenberg Computer Science Department, 4212 Wean Hall, Carnegie Mellon University, Pittsburgh, Pennsylvania Computer Science Department, 4212 Wean Hall, Carnegie Mellon University, Pittsburgh, PennsylvaniaView Profile Authors Info & Claims CHI '86: Proceedings of the SIGCHI Conference on Human Factors in Computing SystemsApril 1986 Pages 41–43https://doi.org/10.1145/22627.22346Published:01 April 1986 4citation366DownloadsMetricsTotal Citations4Total Downloads366Last 12 Months12Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access William Buxton, Roger B. Dannenberg |
CHI | 2 |
| 1985 | Protection for Communication and Sharing in A Personal Computer Network
Roger B. Dannenberg |
ICDCS | 1 |
| 1985 | A Butler Process for Resource Sharing on Spice MachinesabstractA network of personal computers may contain a large amount of distributed computing resources. For a number of reasons it is desirable to share these resources, but sharing is complicated by issues of security and autonomy. A process known as the Butler addresses these problems and provides support for resource sharing. The Butler relies upon a capability-based accounting system called the Banker to monitor the use of local resources. Roger B. Dannenberg, Peter G. Hibbard |
ACM Trans. Inf. Syst. | 1 |
| 1982 | Formal Program Verification Using Symbolic ExecutionabstractSymbolic execution provides a mechanism for formally proving programs correct. A notation is introduced which allows a concise presentation of rules of inference based on symbolic execution. Using this notation, rules of inference are developed to handle a number of language features, including loops and procedures with multiple exits. An attribute grammar is used to formally describe symbolic expression evaluation, and the treatment of function calls with side effects is shown to be straightforward. Because symbolic execution is related to program interpretation, it is an easy-to-comprehend, yet powerful technique. The rules of inference are useful in expressing the semantics of a language and form the basis of a mechanical verification condition generator. Roger B. Dannenberg, George W. Ernst |
IEEE Trans. Software Eng. | 1 |