Mark Cartwright

dblp:81/10919 · also Mark Brozier Cartwright · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-5908-390XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 11 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Like, Comment & Caption: A Decade of Social Media Video Caption Research (2015-2025)
abstract
As video has become the dominant mode of content on platforms such as YouTube, TikTok, and Instagram, captioning has emerged as a critical factor for accessibility, engagement, and visibility. While prior studies have examined different types of social media video captions or communities’ captioning usage, a systematic synthesis has not been undertaken, leading to the risk of proposing interventions that overlook core platform constraints or miss critical accessibility needs. This paper reviews 36 peer-reviewed papers published between 2015 and 2025 across fields such as Human-Computer Interaction (HCI), accessibility, media studies, education, and language learning. We note that captions operate as collective infrastructure co-produced by viewers, creators, and platforms. Deaf and Hard of Hearing (DHH), neurodivergent, and multilingual viewers depend on captions and increasingly expect mechanisms for feedback, while creators face inadequate tool support. Building on these insights, we propose the framework of Participatory Captioning and suggest design implications, highlighting future directions for social media video caption research.
Huong Nguyen, Emma McDonnell, Lloyd May, Alexander Druzenko, Zoobia Saifullah Syeda, Mark Cartwright, Sooyeon Lee
CHI6
2025 Compositional Audio Representation Learning
abstract
Human auditory perception is compositional in nature — we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not disentangle the constituent sound sources. In this work, we learn source-centric audio representations where each sound source is represented using a distinct, disentangled source embedding in the audio representation. We propose two novel approaches to learning source-centric audio representations: a supervised model guided by classification and an unsupervised model guided by feature reconstruction, both of which outperform the baselines. We thoroughly evaluate the design choices of both approaches using an audio classification task. We find that supervision is beneficial to learn source-centric representations, and that reconstructing audio features is more useful than reconstructing spectrograms to learn unsupervised source-centric representations. Leveraging source-centric models can help unlock the potential of greater interpretability and more flexible decoding in machine listening.
Sripathi Sridhar, Mark Cartwright
ICASSP2
2024 Towards a Rich Format for Closed-Captioning
abstract
Closed-captioning is an essential part of viewing audio-visual content for many people, including those who are D/deaf and Hard-of-Hearing. Traditional closed-captioning systems generally consist of a single track of timed text that offers limited options for personalization. Research into extending the capabilities of captioning, such as affective, poetic, and customizable captions has shown a desire among a subset of users for these features, but only in specific contexts. However, due to the difficulty in creating custom stimuli videos utilizing the custom captioning system, comparisons between systems and longitudinal studies have not been pursued. This demo paper introduces Rich Captions, a structured system that allows for a single closed-caption file to be tagged with additional information that can then be flexibly leveraged to render different customizable, creative, and poetic captions from the same file. Additionally, we introduce the Rich Caption Editor 1, a free, open-source software system designed to author, edit, and render rich captions. The system design was informed by a formative design workshop with closed-captioning researchers and advocates. The current design allows researchers to generate reproducible stimuli for closed-captioning studies. Once the design space and user preferences are better understood, the rich captioning framework could be refined to serve a general audience.
Lloyd May, Alex C. Williams, Saad Hassan, Mark Cartwright, Sooyeon Lee
ASSETS4
2024 Unspoken Sound: Identifying Trends in Non-Speech Audio Captioning on YouTube
abstract
High-quality closed captioning of both speech and non-speech elements (e.g., music, sound effects, manner of speaking, and speaker identification) is essential for the accessibility of video content, especially for d/Deaf and hard-of-hearing individuals. While many regions have regulations mandating captioning for television and movies, a regulatory gap remains for the vast amount of web-based video content, including the staggering 500+ hours uploaded to YouTube every minute. Advances in automatic speech recognition have bolstered the presence of captions on YouTube. However, the technology has notable limitations, including the omission of many non-speech elements, which are often crucial for understanding content narratives. This paper examines the contemporary and historical state of non-speech information (NSI) captioning on YouTube through the creation and exploratory analysis of a dataset of over 715k videos. We identify factors that influence NSI caption practices and suggest avenues for future research to enhance the accessibility of online video content.
Lloyd May, Keita Ohshiro, Khang Dang, Sripathi Sridhar, Jhanvi Pai, Magdalena Fuentes, Sooyeon Lee, Mark Cartwright
CHI8
2024 Audio Engineering by People Who Are deaf and Hard of Hearing: Balancing Confidence and Limitations
abstract
With technological advancements, audio engineering has evolved from a domain exclusive to professionals to one open to amateurs. However, research is limited on the accessibility of audio engineering, particularly for deaf, Deaf, and hard of hearing (DHH) individuals. To bridge this gap, we interviewed eight deaf and hard of hearing (dHH) audio engineers in music to understand accessibility in audio engineering. We found that their hearing magnified challenges in audio engineering: insecurities in sound perception undermined their confidence, and the required extra “hearing work” added complexity. As workarounds, participants employed various technologies and techniques, relied on the support of hearing peers, and developed strategies for learning and growth. Through these practices, they navigate audio engineering while balancing confidence and limitations. For future directions, we recommend exploring technologies that reduce insecurities and “hearing work” to empower DHH audio engineers and working toward a DHH-community-driven approach to accessible audio engineering.
Keita Ohshiro, Mark Cartwright
CHI2
2023 Does a Quieter City Mean Fewer Complaints? The Sounds of New York City During Covid-19 Lockdown
abstract
The COVID-19 pandemic had an unprecedented effect in human activity and city landscapes. A very notorious transformation during this period was the change in noise levels and patterns across cities. Small scale studies have show this change in noise levels across different locations in the globe. In this work, we extend these studies by using historical audio data from the SONYC sensor network deployed in New York City. We exploit machine listening models to understand not only noise levels but also patterns, by performing a sound source presence analysis. Finally, we contrast our finding from the acoustic data with noise complaints to better understand the relationship between noise and our perception of it.
Mark Cartwright, Magdalena Fuentes, Charlie Mydlarz, Fabio Miranda 0001, Juan Pablo Bello
ICASSP1
2022 How people who are deaf, Deaf, and hard of hearing use technology in creative sound activities
abstract
Creative sound activities, such as music playing and audio engineering, are said to have been democratized with the development of technology. Yet, the use of technology in creative sound activities by people who are deaf, Deaf, and hard of hearing (DHH) has been underexplored by the research community. To address this gap, we conducted an online survey with 50 DHH participants to understand their use of technology and barriers they face in their creative sound activities. We find DHH people use four types of technology — hearing devices, sound manipulation, sound visualization, and speech-to-text — for three purposes — to improve sound perception via auditory and visual means, to avoid hearing fatigue, and to better communicate with hearing people. We also find their barriers to technology: unknown availability, limited options, and limitations that technology can solve. We discuss opportunities for more inclusive design specific to DHH people’s creative sound activities, as well as facilitating access to information about technology.
Keita Ohshiro, Mark Cartwright
ASSETS2
2022 Active Few-Shot Learning for Sound Event Detection
Yu Wang 0105, Mark Cartwright, Juan Pablo Bello
INTERSPEECH2
2022 Urban Rhapsody: Large-scale exploration of urban soundscapes
abstract
Abstract Noise is one of the primary quality‐of‐life issues in urban environments. In addition to annoyance, noise negatively impacts public health and educational performance. While low‐cost sensors can be deployed to monitor ambient noise levels at high temporal resolutions, the amount of data they produce and the complexity of these data pose significant analytical challenges. One way to address these challenges is through machine listening techniques, which are used to extract features in attempts to classify the source of noise and understand temporal patterns of a city's noise situation. However, the overwhelming number of noise sources in the urban environment and the scarcity of labeled data makes it nearly impossible to create classification models with large enough vocabularies that capture the true dynamism of urban soundscapes. In this paper, we first identify a set of requirements in the yet unexplored domain of urban soundscape exploration. To satisfy the requirements and tackle the identified challenges, we propose Urban Rhapsody, a framework that combines state‐of‐the‐art audio representation, machine learning and visual analytics to allow users to interactively create classification models, understand noise patterns of a city, and quickly retrieve and label audio excerpts in order to create a large high‐precision annotated database of urban sound recordings. We demonstrate the tool's utility through case studies performed by domain experts using data generated over the five‐year deployment of a one‐of‐a‐kind sensor network in New York City.
João Rulff, Fabio Miranda 0001, Marcos Lage, Mark Cartwright, Graham Dove, Juan Pablo Bello, Cláudio T. Silva
Comput. Graph. Forum5
2022 Eliciting Confidence for Improving Crowdsourced Audio Annotations
abstract
In this work we explore confidence elicitation methods for crowdsourcing "soft" labels, e.g., probability estimates, to reduce the annotation costs for domains with ambiguous data. Machine learning research has shown that such "soft" labels are more informative and can reduce the data requirements when training supervised machine learning models. By reducing the number of required labels, we can reduce the costs of slow annotation processes such as audio annotation. In our experiments we evaluated three confidence elicitation methods: 1) "No Confidence" elicitation, 2) "Simple Confidence" elicitation, and 3) "Betting" mechanism for confidence elicitation, at both individual (i.e., per participant) and aggregate (i.e., crowd) levels. In addition, we evaluated the interaction between confidence elicitation methods, annotation types (binary, probability, and z-score derived probability), and "soft" versus "hard" (i.e., binarized) aggregate labels. Our results show that both confidence elicitation mechanisms result in higher annotation quality than the "No Confidence" mechanism for binary annotations at both participant and recording levels. In addition, when aggregating labels at the recording level, results indicate that we can achieve comparable results to those with 10-participant aggregate annotations using fewer annotators if we aggregate "soft" labels instead of "hard" labels. These results suggest that for binary audio annotation using a confidence elicitation mechanism and aggregating continuous labels we can obtain higher annotation quality, more informative labels, with quality differences more pronounced with fewer participants. Finally, we propose a way of integrating these confidence elicitation methods into a two-stage, multi-label annotation pipeline.
Ana Elisa Méndez Méndez, Mark Cartwright, Juan Pablo Bello, Oded Nov
Proc. ACM Hum. Comput. Interact.2
2021 Few-Shot Continual Learning for Audio Classification
abstract
Supervised learning for audio classification typically imposes a fixed class vocabulary, which can be limiting for real-world applications where the target class vocabulary is not known a priori or changes dynamically. In this work, we introduce a few-shot continual learning framework for audio classification, where we can continuously expand a trained base classifier to recognize novel classes based on only few labeled data at inference time. This enables fast and interactive model updates by end-users with minimal human effort. To do so, we leverage the dynamic few-shot learning technique and adapt it to a challenging multi-label audio classification scenario. We incorporate a recent state-of-the-art audio feature extraction model as a backbone and perform a comparative analysis of our approach on two popular audio datasets (ESC-50 and AudioSet). We conduct an in-depth evaluation to illustrate the complexities of the problem and show that, while there is still room for improvement, our method outperforms three baselines on novel class detection while maintaining its performance on base classes.
Yu Wang 0105, Nicholas J. Bryan, Mark Cartwright, Juan Pablo Bello, Justin Salamon
ICASSP3
2021 Specialized Embedding Approximation for Edge Intelligence: A Case Study in Urban Sound Classification
abstract
Embedding models that encode semantic information into low-dimensional vector representations are useful in various machine learning tasks with limited training data. However, these models are typically too large to support inference in small edge devices, which motivates training of smaller yet comparably predictive student embedding models through knowledge distillation (KD). While knowledge distillation traditionally uses the teacher’s original training dataset to train the student, we hypothesize that using a dataset similar to the student’s target domain allows for better compression and training efficiency for the said domain, at the cost of reduced generality across other (non-pertinent) domains. Hence, we introduce Specialized Embedding Approximation (SEA) to train a student featurizer to approximate the teacher’s embedding manifold for a given target domain. We demonstrate the feasibility of SEA in the context of acoustic event classification for urban noise monitoring and show that leveraging a dataset related to this target domain not only improves the baseline performance of the original embedding model but also yields competitive students with >1 order of magnitude lesser storage and activation memory. We further investigate the impact of using random and informed sampling techniques for dimensionality reduction in SEA.
Sangeeta Srivastava, Dhrubojyoti Roy, Mark Cartwright, Juan Pablo Bello, Anish Arora
ICASSP3
2019 Crowdsourcing Multi-label Audio Annotation Tasks with Citizen Scientists
abstract
Annotating rich audio data is an essential aspect of training and evaluating machine listening systems. We approach this task in the context of temporally-complex urban soundscapes, which require multiple labels to identify overlapping sound sources. Typically this work is crowdsourced, and previous studies have shown that workers can quickly label audio with binary annotation for single classes. However, this approach can be difficult to scale when multiple passes with different focus classes are required to annotate data with multiple labels. In citizen science, where tasks are often image-based, annotation efforts typically label multiple classes simultaneously in a single pass. This paper describes our data collection on the Zooniverse citizen science platform, comparing the efficiencies of different audio annotation strategies. We compared multiple-pass binary annotation, single-pass multi-label annotation, and a hybrid approach: hierarchical multi-pass multi-label annotation. We discuss our findings, which support using multi-label annotation, with reference to volunteer citizen scientists' motivations.
Mark Cartwright, Graham Dove, Ana Elisa Méndez Méndez, Juan Pablo Bello, Oded Nov
CHI1
2019 Active Learning for Efficient Audio Annotation and Classification with a Large Amount of Unlabeled Data
abstract
There are many sound classification problems that have target classes which are rare or unique to the context of the problem. For these problems, existing data sets are not sufficient and we must create new problem-specific datasets to train classification models. However, annotating a new dataset for every new problem is costly. Active learning could potentially reduce this annotation cost, but it has been understudied in the context of audio annotation. In this work, we investigate active learning to reduce the annotation cost of a sound classification dataset unique to a particular problem. We evaluate three certainty-based active learning query strategies and propose a new strategy: alternating confidence sampling. Using this strategy, we demonstrate reduced annotation costs when actively training models with both experts and non-experts, and we perform a qualitative analysis on 20k unlabeled recordings to show our approach results in a model that generalizes well to unseen data.
Yu Wang 0105, Ana Elisa Méndez Méndez, Mark Cartwright, Juan Pablo Bello
ICASSP3
2019 Per-Channel Energy Normalization: Why and How
abstract
In the context of automatic speech recognition and acoustic event detection, an adaptive procedure named per-channel energy normalization (PCEN) has recently shown to outperform the pointwise logarithm of mel-frequency spectrogram (logmelspec) as an acoustic frontend. This letter investigates the adequacy of PCEN for spectrogram-based pattern recognition in far-field noisy recordings, both from theoretical and practical standpoints. First, we apply PCEN on various datasets of natural acoustic environments and find empirically that it Gaussianizes distributions of magnitudes while decorrelating frequency bands. Second, we describe the asymptotic regimes of each component in PCEN: temporal integration, gain control, and dynamic range compression. Third, we give practical advice for adapting PCEN parameters to the temporal properties of the noise to be mitigated, the signal to be enhanced, and the choice of time-frequency representation. As it converts a large class of real-world soundscapes into additive white Gaussian noise, PCEN is a computationally efficient frontend for robust detection and classification of acoustic events in heterogeneous environments.
Vincent Lostanlen, Justin Salamon, Mark Cartwright, Brian McFee, Andrew Farnsworth, Steve Kelling, Juan Pablo Bello
IEEE Signal Process. Lett.3
2018 Crowdsourced Pairwise-Comparison for Source Separation Evaluation
abstract
Automated objective methods of audio source separation evaluation are fast, cheap, and require little effort by the investigator. However, their output often correlates poorly with human quality assessments and typically require ground-truth (perfectly separated) signals to evaluate algorithm performance. Subjective multi-stimulus human ratings (e.g. MUSHRA) of audio quality are the gold standard for many tasks, but they are slow and require a great deal of effort to recruit participants and run listening tests. Recent work has shown that a crowdsourced multi-stimulus listening test can have results comparable to lab-based multi-stimulus tests. While these results are encouraging, MUSHRA multi-stimulus tests are limited to evaluating 12 or fewer stimuli, and they require ground-truth stimuli for reference. In this work, we evaluate a web-based pairwise-comparison listening approach that promises to speed and facilitate conducting listening tests, while also addressing some of the shortcomings of multi-stimulus tests. Using audio source separation quality as our evaluation task, we compare our web-based pairwise-comparison listening test to both web-based and lab-based multi-stimulus tests. We find that pairwise-comparison listening tests perform comparably to multi-stimulus tests, but without many of their shortcomings.
Mark Cartwright, Bryan Pardo, Gautham J. Mysore
ICASSP1
2018 Investigating the Effect of Sound-Event Loudness on Crowdsourced Audio Annotations
abstract
Audio annotation is an important step in developing machine-listening systems. It is also a time consuming process, which has motivated investigators to crowdsource audio annotations. However, there are many factors that affect annotations, many of which have not been adequately investigated. In previous work, we investigated the effects of visualization aids and sound scene complexity on the quality of crowdsourced sound-event annotations. In this paper, we extend that work by investigating the effect of sound-event loudness on both sound-event source annotations and sound-event proximity annotations. We find that the sound class, loudness, and annotator bias affect how listeners annotate proximity. We also find that loudness affects recall more than precision and that the strengths of these effects are strongly influenced by the sound class. These findings are not only important for designing effective audio annotation processes, but also for effectively training and evaluating machine-listening systems.
Mark Cartwright, Justin Salamon, Ayanna Seals, Oded Nov, Juan Pablo Bello
ICASSP1
2017 Seeing Sound: Investigating the Effects of Visualizations and Complexity on Crowdsourced Audio Annotations
abstract
Audio annotation is key to developing machine-listening systems; yet, effective ways to accurately and rapidly obtain crowdsourced audio annotations is understudied. In this work, we seek to quantify the reliability/redundancy trade-off in crowdsourced soundscape annotation, investigate how visualizations affect accuracy and efficiency, and characterize how performance varies as a function of audio characteristics. Using a controlled experiment, we varied sound visualizations and the complexity of soundscapes presented to human annotators. Results show that more complex audio scenes result in lower annotator agreement, and spectrogram visualizations are superior in producing higher quality annotations at lower cost of time and human labor. We also found recall is more affected than precision by soundscape complexity, and mistakes can be often attributed to certain sound event characteristics. These findings have implications not only for how we should design annotation tasks and interfaces for audio data, but also how we train and evaluate machine-listening systems.
Mark Cartwright, Ayanna Seals, Justin Salamon, Alex C. Williams, Stefanie Mikloska, Duncan MacConnell, Edith Law, Juan Pablo Bello, Oded Nov
Proc. ACM Hum. Comput. Interact.1
2016 An Approach to Audio-Only Editing for Visually Impaired Seniors
abstract
Older adults and people with vision impairments are increasingly using phones to receive audio-based information and want to publish content online but must use complex audio recording/editing tools that often rely on inaccessible graphical interfaces. This poster describes the design of an accessible audio-based interface for post-processing audio content created by visually impaired seniors. We conducted a diary study with five older adults with vision impairments to understand how to design a system that would allow them to edit content they record using an audio-only interface. Our findings can help inform the development of accessible audio-editing interfaces for people with vision impairments more broadly.
Robin Brewer, Mark Cartwright, Aaron Karp, Bryan Pardo, Anne Marie Piper
ASSETS2
2016 Fast and easy crowdsourced perceptual audio evaluation
abstract
Automated objective methods of audio evaluation are fast, cheap, and require little effort by the investigator. However, objective evaluation methods do not exist for the output of all audio processing algorithms, often have output that correlates poorly with human quality assessments, and require ground truth data in their calculation. Subjective human ratings of audio quality are the gold standard for many tasks, but are expensive, slow, and require a great deal of effort to recruit subjects and run listening tests. Moving listening tests from the lab to the micro-task labor market of Amazon Mechanical Turk speeds data collection and reduces investigator effort. However, it also reduces the amount of control investigators have over the testing environment, adding new variability and potential biases to the data. In this work, we compare multiple stimulus listening tests performed in a lab environment to multiple stimulus listening tests performed in web environment on a population drawn from Mechanical Turk.
Mark Cartwright, Bryan Pardo, Gautham J. Mysore, Matthew Hoffman 0001
ICASSP1
2015 VocalSketch: Vocally Imitating Audio Concepts
abstract
A natural way of communicating an audio concept is to imitate it with one's voice. This creates an approximation of the imagined sound (e.g. a particular owl's hoot), much like how a visual sketch approximates a visual concept (e.g a drawing of the owl). If a machine could understand vocal imitations, users could communicate with software in this natural way, enabling new interactions (e.g. programming a music synthesizer by imitating the desired sound with one's voice). In this work, we collect thousands of crowd-sourced vocal imitations of a large set of diverse sounds, along with data on the crowd's ability to correctly label these vocal imitations. The resulting data set will help the research community understand which audio concepts can be effectively communicated with this approach. We have released the data set so the community can study the related issues and build systems that leverage vocal imitation as an interaction modality.
Mark Cartwright, Bryan Pardo
CHI1
2014 MIXPLORATION: rethinking the audio mixer interface
abstract
A typical audio mixer interface consists of faders and knobs that control the amplitude level as well as processing (e.g. equalization, compression and reverberation) parameters of individual tracks. This interface, while widely used and effective for optimizing a mix, may not be the best interface to facilitate exploration of different mixing options. In this work, we rethink the mixer interface, describing an alternative interface for exploring the space of possible mixes of four audio tracks. In a user study with 24 participants, we compared the effectiveness of this interface to the traditional paradigm for exploring alternative mixes. In the study, users responded that the proposed alternative interface facilitated exploration and that they considered the process of rating mixes to be beneficial.
Mark Cartwright, Bryan Pardo, Joshua D. Reiss
IUI1
2014 SynthAssist: an audio synthesizer programmed with vocal imitation
abstract
While programming an audio synthesizer can be difficult, if a user has a general idea of the sound they are trying to program, they may be able to imitate it with their voice. In this technical demonstration, we demonstrate SynthAssist, a system that allows the user to program an audio synthesizer using vocal imitation and interactive feedback. This system treats synthesizer programming as an audio information retrieval task. To account for the limitations of the human voice, it compares vocal imitations to synthesizer sounds by using both absolute and relative temporal shapes of relevant audio features, and it refines the query and feature weights using relevance feedback.
Mark Cartwright, Bryan Pardo
ACM Multimedia1