Charles P. Martin

dblp:208/2413 · also Charles Martin 0001, Charles Patrick Martin · DBLP profile ↗
← Back
15ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-5683-7529ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music Editing
abstract
Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided editing methods rely on pretrained diffusion models by involving forward-backward diffusion processes. However, these methods often struggle to preserve the musical content. Additionally, text instructions alone usually fail to accurately describe the desired music. In this paper, we propose two music editing methods that improve the consistency between the original and edited music by leveraging score distillation. The first method, SteerMusic, is a coarse-grained zero-shot editing approach using delta denoising score. The second method, SteerMusic+, enables fine-grained personalized music editing by manipulating a concept token that represents a user-defined musical style. SteerMusic+ allows for the editing of music into user-defined musical styles that cannot be achieved by the text instructions alone. Experimental results show that our methods outperform existing approaches in preserving both music content consistency and editing fidelity. User studies further validate that our methods achieve superior music editing quality.
Xinlei Niu, Kin Wai Cheuk, Jing Zhang 0052, Naoki Murata, Chieh-Hsin Lai, Michele Mancusi, Woosung Choi, Giorgio Fabbro, Wei-Hsiang Liao 0001, Charles P. Martin, Yuki Mitsufuji
AAAI10
2025 Seeing the Sound: Supporting Musical Collaboration with Augmented Reality
abstract
In musical collaboration, digital musical instruments often hinder effective communication and engagement by restricting visibility and limiting gestural and non-verbal interactions.These challenges reduce musicians' situational awareness and complicate cohesive performance.To address this, we developed a head-mounted augmented reality (AR) system to enhance collaborative musical experiences by visualising musicians' hand movements, eye gaze positions, and instrument interactions in real-time.We conducted a user study involving four pairs of musicians performing live music using different AR interface configurations.The results suggest that the AR system can enhance situational awareness and assist collaboration, as reflected in questionnaire responses.Interviews indicated that real-time visualisations of bodily movements and interactions helped participants better understand the collaborative process and anticipate their collaborators' actions.These findings point to the potential of AR-assisted visualisation to support creative collaboration by tailoring visual information to different needs.Future research could explore its application in broader contexts of real-time creative cooperation.
Mingze Xi, Matt Adcock, Charles P. Martin
Creativity & Cognition4
2024 SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation
abstract
We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently trained under limited computational resources. The integration of a contrastive learning strategy further enhances the connection between text conditions and the generated outputs, resulting in coherent and high-fidelity performance. Our experiments demonstrate that SoundLoCD outperforms the baseline with greatly reduced computational resources. A comprehensive ablation study further validates the contribution of each component within SoundLoCD1.
Xinlei Niu, Jing Zhang 0052, Christian Walder, Charles P. Martin
ICASSP4
2024 Latent Optimal Paths by Gumbel Propagation for Variational Bayesian Dynamic Programming
abstract
We propose the stochastic optimal path which solves the classical optimal path problem by a probability-softening solution. This unified approach transforms a wide range of DP problems into directed acyclic graphs in which all paths follow a Gibbs distribution. We show the equivalence of the Gibbs distribution to a message-passing algorithm by the properties of the Gumbel distribution and give all the ingredients required for variational Bayesian inference of a latent path, namely Bayesian dynamic programming (BDP). We demonstrate the usage of BDP in the latent space of variational autoencoders (VAEs) and propose the BDP-VAE which captures structured sparse optimal paths as latent variables. This enables end-to-end training for generative tasks in which models rely on unobserved structural information. At last, we validate the behavior of our approach and showcase its applicability in two real-world applications: text-to-speech and singing voice synthesis. Our implementation code is available at https://github.com/XinleiNIU/LatentOptimalPathsBayesianDP.
Xinlei Niu, Christian Walder, Jing Zhang 0052, Charles P. Martin
ICML4
2024 HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
abstract
We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with contrastive learning. HybridVC supports text and audio prompts, enabling more flexible voice style conversion. HybridVC models a latent distribution conditioned on speaker embeddings acquired by a pretrained speaker encoder and optimises style text embeddings to align with the speaker style information through contrastive learning in parallel. Therefore, HybridVC can be efficiently trained under limited computational resources. Our experiments demonstrate HybridVC's superior training efficiency and its capability for advanced multimodal voice style conversion. This underscores its potential for widespread applications such as user-defined personalised voice in various social media platforms. A comprehensive ablation study further validates the effectiveness of our method.
Xinlei Niu, Jing Zhang 0052, Charles P. Martin
INTERSPEECH3
2023 Embodying an Interactive AI for Dance Through Movement Ideation
abstract
What expectations exist in the minds of dancers when interacting with a generative machine learning model? During two workshop events, experienced dancers explore these expectations through improvisation and role-play, embodying an imagined AI-dancer. The dancers explored how intuited flow, shared images, and the concept of a human replica might work in their imagined AI-human interaction. Our findings challenge existing assumptions about what is desired from generative models of dance, such as expectations of realism, and how such systems should be evaluated. We further advocate that such models should celebrate non-human artefacts, focus on the potential for serendipitous moments of discovery, and that dance practitioners should be included in their development. Our concrete suggestions show how our findings can be adapted into the development of improved generative and interactive machine learning models for dancers’ creative practice.
Benedikte Wallace, Clarice Hilton, Kristian Nymoen, Jim Tørresen, Charles P. Martin, Rebecca Fiebrink
Creativity & Cognition5
2022 Spatial-Temporal-Class Attention Network for Acoustic Scene Classification
abstract
Acoustic scene classification, where a scene is identified from a sound recording, is a difficult problem that is much less studied than similar problems in computer vision. Re-cent advances in attention-based convolution neural networks (CNNs) can be applied to audio data by operating on two dimensional spectrograms, where frequency and time infor-mation have been separated, rather than a raw audio signal. Typical CNNs have difficulty coping with this problem due to the temporal aspects of acoustic data. In this research we propose a novel and intuitive CNN-based architecture with attention mechanisms called the spatial-temporal-class attention network (STCANet). The STCANet consists of a spatial-temporal attention and a class attention which extracts in-formation along with frequency, temporal, and the class di-mension of spectrograms. In our experiments, the STCANet achieved 75.6%, 95.4%, and 97.0% accuracy on TUT 2018, TAU 2020, and ESC-I0 datasets that are competitive results compared with previous works. Our contributions include this novel network design and a detailed analysis of how attention allows these results to be achieved.
Xinlei Niu, Charles P. Martin
ICME2
2021 Learning Embodied Sound-Motion Mappings: Evaluating AI-Generated Dance Improvisation
abstract
Through dance, a wide range of emotions can be expressed. As virtual agents and robots continue to become part of our daily lives, the need for them to efficiently convey emotion and intent increases. When trained to dance, to what extent can AI learn to model the tacit mappings between sound and motion? Here, we explore the creative capacity of a generative model trained on 3D motion capture recordings of improvised dance. We perform a perceptual judgment experiment wherein respondents rate movement generated by our model as well as human performances. While the sound-motion mappings remain somewhat elusive, particularly when compared to examples of human dance, our study shows that in certain aspects related to perceived dance-likeness and expressivity, the model successfully mimics human dance movement. By employing a perceptual study to evaluate our generative model, we aim to further our ability to understand the affordances and limitations of creative AI.
Benedikte Wallace, Charles P. Martin, Jim Tørresen, Kristian Nymoen
Creativity & Cognition2
2021 Environmental Adaptation of Robot Morphology and Control Through Real-World Evolution
abstract
Robots operating in the real world will experience a range of different environments and tasks. It is essential for the robot to have the ability to adapt to its surroundings to work efficiently in changing conditions. Evolutionary robotics aims to solve this by optimizing both the control and body (morphology) of a robot, allowing adaptation to internal, as well as external factors. Most work in this field has been done in physics simulators, which are relatively simple and not able to replicate the richness of interactions found in the real world. Solutions that rely on the complex interplay among control, body, and environment are therefore rarely found. In this article, we rely solely on real-world evaluations and apply evolutionary search to yield combinations of morphology and control for our mechanically self-reconfiguring quadruped robot. We evolve solutions on two distinct physical surfaces and analyze the results in terms of both control and morphology. We then transition to two previously unseen surfaces to demonstrate the generality of our method. We find that the evolutionary search finds high-performing and diverse morphology-controller configurations by adapting both control and body to the different properties of the physical environments. We additionally find that morphology and control vary with statistical significance between the environments. Moreover, we observe that our method allows for morphology and control parameters to transfer to previously unseen terrains, demonstrating the generality of our approach.
Tønnes F. Nygaard, Charles P. Martin, Gerard David Howard, Jim Tørresen, Kyrre Glette
Evol. Comput.2
2020 Towards Movement Generation with Audio Features
Benedikte Wallace, Charles P. Martin, Jim Tørresen, Kristian Nymoen
ICCC2
2019 Evolving Robots on Easy Mode: Towards a Variable Complexity Controller for Quadrupeds
Tønnes F. Nygaard, Charles P. Martin, Jim Tørresen, Kyrre Glette
EvoApplications2
2019 Self-Modifying Morphology Experiments with DyRET: Dynamic Robot for Embodied Testing
abstract
If robots are to become ubiquitous, they will need to be able to adapt to complex and dynamic environments. Robots that can adapt their bodies while deployed might be flexible and robust enough to meet this challenge. Previous work on dynamic robot morphology has focused on simulation, combining simple modules, or switching between locomotion modes. Here, we present an alternative approach: a self-reconfigurable morphology that allows a single four-legged robot to actively adapt the length of its legs to different environments. We report the design of our robot, as well as the results of a study that verifies the performance impact of self-reconfiguration. This study compares three different control and morphology pairs under different levels of servo supply voltage in the lab. We also performed preliminary tests in different uncontrolled outdoor environments to see if changes to the external environment supports our findings in the lab. Our results show better performance with an adaptable body, lending evidence to the value of self-reconfiguration for quadruped robots.
Tønnes F. Nygaard, Charles P. Martin, Jim Tørresen, Kyrre Glette
ICRA2
2018 Real-world evolution adapts robot morphology and control to hardware limitations
abstract
For robots to handle the numerous factors that can affect them in the real world, they must adapt to changes and unexpected events. Evolutionary robotics tries to solve some of these issues by automatically optimizing a robot for a specific environment. Most of the research in this field, however, uses simplified representations of the robotic system in software simulations. The large gap between performance in simulation and the real world makes it challenging to transfer the resulting robots to the real world. In this paper, we apply real world multi-objective evolutionary optimization to optimize both control and morphology of a four-legged mammal-inspired robot. We change the supply voltage of the system, reducing the available torque and speed of all joints, and study how this affects both the fitness, as well as the morphology and control of the solutions. In addition to demonstrating that this real-world evolutionary scheme for morphology and control is indeed feasible with relatively few evaluations, we show that evolution under the different hardware limitations results in comparable performance for low and moderate speeds, and that the search achieves this by adapting both the control and the morphology of the robot.
Tønnes F. Nygaard, Charles P. Martin, Eivind Samuelsen, Jim Tørresen, Kyrre Glette
GECCO2
2016 Intelligent Agents and Networked Buttons Improve Free-Improvised Ensemble Music-Making on Touch-Screens
abstract
We present the results of two controlled studies of free-improvised ensemble music-making on touch-screens. In our system, updates to an interface of harmonically-selected pitches are broadcast to every touch-screen in response to either a performer pressing a GUI button, or to interventions from an intelligent agent. In our first study, analysis of survey results and performance data indicated significant effects of the button on performer preference, but of the agent on performance length. In the second follow-up study, a mixed-initiative interface, where the presence of the button was interlaced with agent interventions, was developed to leverage both approaches. Comparison of this mixed-initiative interface with the always-on button-plus-agent condition of the first study demonstrated significant preferences for the former. The different approaches were found to shape the creative interactions that take place. Overall, this research offers evidence that an intelligent agent and a networked GUI both improve aspects of improvised ensemble music-making.
Charles P. Martin, Henry J. Gardner, Ben Swift, Michael A. Martin
CHI1
2014 Exploring percussive gesture on iPads with ensemble metatone
abstract
Percussionists are unique among western classical instrumentalists in that their artistic practice is defined by an approach to interaction rather than their instruments. While percussionists are accustomed to exploring non-traditional objects to create music, these objects have yet to encompass touch-screen computing devices to any great extent. The proliferation and popularity of these devices now presents an opportunity to explore their use in combining computer-generated sound together with percussive interaction in a musical ensemble.
Charles P. Martin, Henry J. Gardner, Ben Swift
CHI1