VLDB 2026 Research / reviewers in the wild / expert
Charles P. Martin
dblp:208/2413 · also Charles Martin 0001, Charles Patrick Martin
· DBLP profile ↗
15ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-5683-7529ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music EditingabstractMusic editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods often struggle to preserve the musical content. Additionally, text instructions alone usually fail to accurately describe the desired music. In this paper, we propose two music editing methods that improve the consistency between the original and edited music by leveraging score distillation. The first method, SteerMusic, is a coarse-grained zero-shot editing approach using delta denoising score. The second method, SteerMusic+, enables fine-grained personalized music editing by manipulating a concept token that represents a user-defined musical style. SteerMusic+ allows for the editing of music into user-defined musical styles that cannot be achieved by the text instructions alone. Experimental results show that our methods outperform existing approaches in preserving both music content consistency and editing fidelity. User studies further validate that our methods achieve superior music editing quality. Xinlei Niu, Kin Wai Cheuk, Jing Zhang 0052, Naoki Murata, Chieh-Hsin Lai, Michele Mancusi, Woosung Choi, Giorgio Fabbro, Wei-Hsiang Liao 0001, Charles P. Martin, Yuki Mitsufuji |
AAAI | 10 |
| 2025 | Seeing the Sound: Supporting Musical Collaboration with Augmented RealityabstractIn musical collaboration, digital musical instruments often hinder effective communication and engagement by restricting visibility and limiting gestural and non-verbal interactions.These challenges reduce musicians' situational awareness and complicate cohesive performance.To address this, we developed a head-mounted augmented reality (AR) system to enhance collaborative musical experiences by visualising musicians' hand movements, eye gaze positions, and instrument interactions in real-time.We conducted a user study involving four pairs of musicians performing live music using different AR interface configurations.The results suggest that the AR system can enhance situational awareness and assist collaboration, as reflected in questionnaire responses.Interviews indicated that real-time visualisations of bodily movements and interactions helped participants better understand the collaborative process and anticipate their collaborators' actions.These findings point to the potential of AR-assisted visualisation to support creative collaboration by tailoring visual information to different needs.Future research could explore its application in broader contexts of real-time creative cooperation. Mingze Xi, Matt Adcock, Charles P. Martin |
Creativity & Cognition | 4 |
| 2024 | SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound GenerationabstractWe present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently trained under limited computational resources. The integration of a contrastive learning strategy further enhances the connection between text conditions and the generated outputs, resulting in coherent and high-fidelity performance. Our experiments demonstrate that SoundLoCD outperforms the baseline with greatly reduced computational resources. A comprehensive ablation study further validates the contribution of each component within SoundLoCD1. Xinlei Niu, Jing Zhang 0052, Christian Walder, Charles P. Martin |
ICASSP | 4 |
| 2024 | Latent Optimal Paths by Gumbel Propagation for Variational Bayesian Dynamic ProgrammingabstractWe propose the stochastic optimal path which solves the classical optimal path problem by a probability-softening solution. This unified approach transforms a wide range of DP problems into directed acyclic graphs in which all paths follow a Gibbs distribution. We show the equivalence of the Gibbs distribution to a message-passing algorithm by the properties of the Gumbel distribution and give all the ingredients required for variational Bayesian inference of a latent path, namely Bayesian dynamic programming (BDP). We demonstrate the usage of BDP in the latent space of variational autoencoders (VAEs) and propose the BDP-VAE which captures structured sparse optimal paths as latent variables. This enables end-to-end training for generative tasks in which models rely on unobserved structural information. At last, we validate the behavior of our approach and showcase its applicability in two real-world applications: text-to-speech and singing voice synthesis. Our implementation code is available at https://github.com/XinleiNIU/LatentOptimalPathsBayesianDP. Xinlei Niu, Christian Walder, Jing Zhang 0052, Charles P. Martin |
ICML | 4 |
| 2024 | HybridVC: Efficient Voice Style Conversion with Text and Audio PromptsabstractWe introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with contrastive learning. HybridVC supports text and audio prompts, enabling more flexible voice style conversion. HybridVC models a latent distribution conditioned on speaker embeddings acquired by a pretrained speaker encoder and optimises style text embeddings to align with the speaker style information through contrastive learning in parallel. Therefore, HybridVC can be efficiently trained under limited computational resources. Our experiments demonstrate HybridVC's superior training efficiency and its capability for advanced multimodal voice style conversion. This underscores its potential for widespread applications such as user-defined personalised voice in various social media platforms. A comprehensive ablation study further validates the effectiveness of our method. Xinlei Niu, Jing Zhang 0052, Charles P. Martin |
INTERSPEECH | 3 |
| 2023 | Embodying an Interactive AI for Dance Through Movement IdeationabstractWhat expectations exist in the minds of dancers when interacting with a generative machine learning model? During two workshop events, experienced dancers explore these expectations through improvisation and role-play, embodying an imagined AI-dancer. The dancers explored how intuited flow, shared images, and the concept of a human replica might work in their imagined AI-human interaction. Our findings challenge existing assumptions about what is desired from generative models of dance, such as expectations of realism, and how such systems should be evaluated. We further advocate that such models should celebrate non-human artefacts, focus on the potential for serendipitous moments of discovery, and that dance practitioners should be included in their development. Our concrete suggestions show how our findings can be adapted into the development of improved generative and interactive machine learning models for dancers’ creative practice. Benedikte Wallace, Clarice Hilton, Kristian Nymoen, Jim Tørresen, Charles P. Martin, Rebecca Fiebrink |
Creativity & Cognition | 5 |
| 2022 | Spatial-Temporal-Class Attention Network for Acoustic Scene ClassificationabstractAcoustic scene classification, where a scene is identified from a sound recording, is a difficult problem that is much less studied than similar problems in computer vision. Re-cent advances in attention-based convolution neural networks (CNNs) can be applied to audio data by operating on two dimensional spectrograms, where frequency and time infor-mation have been separated, rather than a raw audio signal. Typical CNNs have difficulty coping with this problem due to the temporal aspects of acoustic data. In this research we propose a novel and intuitive CNN-based architecture with attention mechanisms called the spatial-temporal-class attention network (STCANet). The STCANet consists of a spatial-temporal attention and a class attention which extracts in-formation along with frequency, temporal, and the class di-mension of spectrograms. In our experiments, the STCANet achieved 75.6%, 95.4%, and 97.0% accuracy on TUT 2018, TAU 2020, and ESC-I0 datasets that are competitive results compared with previous works. Our contributions include this novel network design and a detailed analysis of how attention allows these results to be achieved. Xinlei Niu, Charles P. Martin |
ICME | 2 |
| 2021 | Learning Embodied Sound-Motion Mappings: Evaluating AI-Generated Dance ImprovisationabstractThrough dance, a wide range of emotions can be expressed. As virtual agents and robots continue to become part of our daily lives, the need for them to efficiently convey emotion and intent increases. When trained to dance, to what extent can AI learn to model the tacit mappings between sound and motion? Here, we explore the creative capacity of a generative model trained on 3D motion capture recordings of improvised dance. We perform a perceptual judgment experiment wherein respondents rate movement generated by our model as well as human performances. While the sound-motion mappings remain somewhat elusive, particularly when compared to examples of human dance, our study shows that in certain aspects related to perceived dance-likeness and expressivity, the model successfully mimics human dance movement. By employing a perceptual study to evaluate our generative model, we aim to further our ability to understand the affordances and limitations of creative AI. Benedikte Wallace, Charles P. Martin, Jim Tørresen, Kristian Nymoen |
Creativity & Cognition | 2 |
| 2021 | Environmental Adaptation of Robot Morphology and Control Through Real-World EvolutionabstractRobots operating in the real world will experience a range of different environments and tasks. It is essential for the robot to have the ability to adapt to its surroundings to work efficiently in changing conditions. Evolutionary robotics aims to solve this by optimizing both the control and body (morphology) of a robot, allowing adaptation to internal, as well as external factors. Most work in this field has been done in physics simulators, which are relatively simple and not able to replicate the richness of interactions found in the real world. Solutions that rely on the complex interplay among control, body, and environment are therefore rarely found. In this article, we rely solely on real-world evaluations and apply evolutionary search to yield combinations of morphology and control for our mechanically self-reconfiguring quadruped robot. We evolve solutions on two distinct physical surfaces and analyze the results in terms of both control and morphology. We then transition to two previously unseen surfaces to demonstrate the generality of our method. We find that the evolutionary search finds high-performing and diverse morphology-controller configurations by adapting both control and body to the different properties of the physical environments. We additionally find that morphology and control vary with statistical significance between the environments. Moreover, we observe that our method allows for morphology and control parameters to transfer to previously unseen terrains, demonstrating the generality of our approach. Tønnes F. Nygaard, Charles P. Martin, Gerard David Howard, Jim Tørresen, Kyrre Glette |
Evol. Comput. | 2 |
| 2020 | Towards Movement Generation with Audio Features
Benedikte Wallace, Charles P. Martin, Jim Tørresen, Kristian Nymoen |
ICCC | 2 |
| 2019 | Evolving Robots on Easy Mode: Towards a Variable Complexity Controller for Quadrupeds
Tønnes F. Nygaard, Charles P. Martin, Jim Tørresen, Kyrre Glette |
EvoApplications | 2 |
| 2019 | Self-Modifying Morphology Experiments with DyRET: Dynamic Robot for Embodied TestingabstractIf robots are to become ubiquitous, they will need to be able to adapt to complex and dynamic environments. Robots that can adapt their bodies while deployed might be flexible and robust enough to meet this challenge. Previous work on dynamic robot morphology has focused on simulation, combining simple modules, or switching between locomotion modes. Here, we present an alternative approach: a self-reconfigurable morphology that allows a single four-legged robot to actively adapt the length of its legs to different environments. We report the design of our robot, as well as the results of a study that verifies the performance impact of self-reconfiguration. This study compares three different control and morphology pairs under different levels of servo supply voltage in the lab. We also performed preliminary tests in different uncontrolled outdoor environments to see if changes to the external environment supports our findings in the lab. Our results show better performance with an adaptable body, lending evidence to the value of self-reconfiguration for quadruped robots. Tønnes F. Nygaard, Charles P. Martin, Jim Tørresen, Kyrre Glette |
ICRA | 2 |
| 2018 | Real-world evolution adapts robot morphology and control to hardware limitationsabstractFor robots to handle the numerous factors that can affect them in the real world, they must adapt to changes and unexpected events. Evolutionary robotics tries to solve some of these issues by automatically optimizing a robot for a specific environment. Most of the research in this field, however, uses simplified representations of the robotic system in software simulations. The large gap between performance in simulation and the real world makes it challenging to transfer the resulting robots to the real world. In this paper, we apply real world multi-objective evolutionary optimization to optimize both control and morphology of a four-legged mammal-inspired robot. We change the supply voltage of the system, reducing the available torque and speed of all joints, and study how this affects both the fitness, as well as the morphology and control of the solutions. In addition to demonstrating that this real-world evolutionary scheme for morphology and control is indeed feasible with relatively few evaluations, we show that evolution under the different hardware limitations results in comparable performance for low and moderate speeds, and that the search achieves this by adapting both the control and the morphology of the robot. Tønnes F. Nygaard, Charles P. Martin, Eivind Samuelsen, Jim Tørresen, Kyrre Glette |
GECCO | 2 |
| 2016 | Intelligent Agents and Networked Buttons Improve Free-Improvised Ensemble Music-Making on Touch-ScreensabstractWe present the results of two controlled studies of free-improvised ensemble music-making on touch-screens. In our system, updates to an interface of harmonically-selected pitches are broadcast to every touch-screen in response to either a performer pressing a GUI button, or to interventions from an intelligent agent. In our first study, analysis of survey results and performance data indicated significant effects of the button on performer preference, but of the agent on performance length. In the second follow-up study, a mixed-initiative interface, where the presence of the button was interlaced with agent interventions, was developed to leverage both approaches. Comparison of this mixed-initiative interface with the always-on button-plus-agent condition of the first study demonstrated significant preferences for the former. The different approaches were found to shape the creative interactions that take place. Overall, this research offers evidence that an intelligent agent and a networked GUI both improve aspects of improvised ensemble music-making. Charles P. Martin, Henry J. Gardner, Ben Swift, Michael A. Martin |
CHI | 1 |
| 2014 | Exploring percussive gesture on iPads with ensemble metatoneabstractPercussionists are unique among western classical instrumentalists in that their artistic practice is defined by an approach to interaction rather than their instruments. While percussionists are accustomed to exploring non-traditional objects to create music, these objects have yet to encompass touch-screen computing devices to any great extent. The proliferation and popularity of these devices now presents an opportunity to explore their use in combining computer-generated sound together with percussive interaction in a musical ensemble. Charles P. Martin, Henry J. Gardner, Ben Swift |
CHI | 1 |