VLDB 2026 Research / reviewers in the wild / expert
Youngmoo E. Kim
dblp:89/1201
· DBLP profile ↗
12ranked-venue papers
0as first author
0since 2021 · last 2020
0009-0005-0922-1787ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6Systems, architecture and hardware · 3Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
3 papers |
Audio and music processing · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Interaction techniques and input · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
music information retrieval |
0.2 | 1 | 2014 | Representing Musical Patterns via the Rhythmic Style Histogram Feature · ACM Multimedia 2014 |
Interaction techniques and input › input sensing
gesture sensing |
0.1 | 1 | 2011 | Multidimensional gesture sensing at the piano keyboard · CHI 2011 |
Interaction techniques and input
musical interface |
0.1 | 1 | 2011 | Multidimensional gesture sensing at the piano keyboard · CHI 2011 |
Audio and music processing › speech enhancement
dereverberation |
0.1 | 1 | 2007 | Blind channel identification for speech dereverberation using l1-norm sparse learning · NIPS 2007 |
Audio and music processing › room acoustics
room impulse response modeling |
0.1 | 1 | 2007 | Blind channel identification for speech dereverberation using l1-norm sparse learning · NIPS 2007 |
Audio and music processing › music analysis
onset detection |
0.1 | 1 | 2014 | Representing Musical Patterns via the Rhythmic Style Histogram Feature · ACM Multimedia 2014 |
Mathematical optimization › regularization
l1-regularized least squares |
0.0 | 1 | 2007 | Blind channel identification for speech dereverberation using l1-norm sparse learning · NIPS 2007 |
Mathematical optimization › least squares
regularized least squares |
0.0 | 1 | 2007 | Blind channel identification for speech dereverberation using l1-norm sparse learning · NIPS 2007 |
Methods — techniques the papers use, named apart from their topics
user study · 0.2unsupervised learning · 0.2supervised learning · 0.2histogram feature · 0.2l1-norm sparse learning · 0.1convex optimization · 0.1bayesian inference · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Toward Accurate Sensing with Knitted Fabric: Applications and Technical ConsiderationsabstractFabric sensors have been introduced to enable flexible touch-based interaction. We advance the technical capabilities of a scalable and low-profile knitted capacitive touch sensing system by introducing methods to improve its touch localization accuracy. The sensor hardware design tends toward minimalism by using a single conductive yarn and two external connections located at each endpoint. Fewer connectors simplify the textile system integration, but this comes at the expense of reduced signal information output from the system. The electrical continuity of the sensing element, essential to the process of knitting, also increases the uncertainty of localizing touch. We propose using Bode analysis to measure changes in signal due to capacitive touch, as well as design a new algorithm, MSD, which retains the most significant aspects of the signal in terms of touch location identification. We do not classify location of touch, but focus on an invariant signal representation. To evaluate our methods, we introduce ELD, a distance metric to compute the similarity of pairs of key-presses, generalizable to computing distances of tensors of varying lengths. Our experiments show that the proposed sensing method results in high-fidelity signals. Furthermore, the sparse representation of key-presses produced by MSD significantly increases separability between different touch locations. Possible applications based on these sensors are also illustrated through prototypes and use case descriptions. Richard Vallett, Denisa Qori McDonald, Geneviève Dion, Youngmoo E. Kim, Ali Shokoufandeh |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2015 | Acoustic Features for Recognizing Musical Artist InfluenceabstractMusicologists have been interested in the topic of influence between composers for years and have developed methods and heuristics for recognizing influence in classical music. While these methods work well for music where the score is the primary source of information, this type of analysis is not well suited for modern popular music where the audio recording itself is arguably the primary representation. This paper presents two audio content-based systems for influence recognition: a system using a spectral representation (Constant-q transform) and support vector machines and another system that obtains features by using a deep belief network and then logistic regression for classification. The system using the spectral representation provides a baseline for future comparisons and evidence to support the idea that influence recognition can be performed using information extracted from the audio signal. The other system attempts to improve performance by using a deep belief network to learn features useful for influence recognition by mapping data extracted from the audio signal to labeled influence data. A dataset of about 77,000 30-second audio clips, consisting of retail previews of popular music tracks was gathered for this work. These songs were chosen from expertly-labeled influence relationship information gathered by the editors of the AllMusic guide. Brandon G. Morton, Youngmoo E. Kim |
ICMLA | 2 |
| 2014 | Rapidly learning musical beats in the presence of environmental and robot ego noiseabstractHumans can often learn high-level features of a piece of music, such as beats, from only a few seconds of audio. If robots could obtain this information just as rapidly, they would be more capable of musical interaction without needing long lead times to learn the music. The presence of robot ego noise, however, makes accurately analyzing music more difficult. In this paper, we focus on the task of learning musical beats, which are often identifiable to humans even in noisy environments such as bars. Learning beats would not only help robots to synchronize their responses to music, but could lead to learning other aspects of musical audio, such as other repeated events, timbrel aspects, and more. We introduce a novel algorithm utilizing stacked spectrograms, in which each column contains frequency bins from multiple instances in time, as well as Probabilistic Latent Component Analysis (PLCA) to learn beats in noisy audio. The stacked spectrograms are exploited to find time-varying spectral characteristics of acoustic components, and PLCA is used to learn and separate the components and find those containing beats. We demonstrate that this system can learn musical beats even when only provided with a few seconds of noisy audio. David Grunberg, Youngmoo E. Kim |
IROS | 2 |
| 2014 | A lightweight, cross-platform, multiuser robot visualization using the cloudabstractCloud robotics emphasizes harnessing the power of the Web for robotics. Modern mobile devices connect to the Web and are convenient user interfaces. We decided to explore what was possible at the intersection of robotics, the Web, and mobile devices by creating a mobile web interface for a humanoid robot. This paper describes our implementation of a monitoring interface for high degree of freedom (DOF) robots that works with both desktop and mobile devices. Using only standard web technologies, our application provides a rich 3D interface that displays the robot's pose, orientation, and sensor data, and can update at 30Hz. It is easy to use, because there is no software for the user to install; it runs using the device's mobile browser. The web interface can be deployed on a private or public cloud, and is designed to scale to support hundreds or thousands of viewers by utilizing cloud services. The system was successfully tested with two different robots and on multiple browsers and mobile devices. William Hilton, Daniel M. Lofaro, Youngmoo E. Kim |
IROS | 3 |
| 2014 | Representing Musical Patterns via the Rhythmic Style Histogram FeatureabstractWhen listening to music, humans often focus on melodic and rhythmic elements to identify specific songs or genres. While these representations may be quite simple, they still capture and differentiate higher level aspects of music such as expressive intent and musical style. In this work we seek to extract and represent rhythmic patterns from a polyphonic corpus of audio encompassing a number of styles. A compact feature is designed that probabilistically models rhythmic activations within musical beat divisions through histograms of Inter-Onset-Intervals (IOI). Onset detection functions are calculated from multiple frequency bands of a perceptually motivated filter bank. This allows for patterns of lower pitched and higher pitched onsets to be described separately. Through a set of supervised and unsupervised experiments, we show that this feature is well suited for a variety of tasks in which quantifying rhythmic style is necessary. Matthew Prockup, Jeffrey J. Scott, Youngmoo E. Kim |
ACM Multimedia | 3 |
| 2013 | Utilizing music technology as a model for creativity development in K-12 educationabstractMany students are highly engaged, motivated, and intellectually stimulated by music outside of the classroom. In 2012, the US ranked 17th among developed countries in education. A major commonality in nations outperforming the US is a deeper focus on the arts. We argue it necessary to find new ways to engage students in music education. In this initial work, we demonstrate that teaching with music technology provides an affordable point of entry for non-trained music students to express their musical sensibilities. Computer-based tools have become the standard for the music industry. We posit that music technology classes serve as an excellent environment for creative development, offering self-awareness of one's creative process, experiential flow learning, and creative thinking skills. David S. Rosen, Erik M. Schmidt, Youngmoo E. Kim |
Creativity & Cognition | 3 |
| 2011 | Multidimensional gesture sensing at the piano keyboardabstractIn this paper we present a new keyboard interface for computer music applications. Where traditional keyboard controllers report the velocity of each key-press, our interface senses up to five separate dimensions: velocity, percussiveness, rigidity, weight, and depth. These dimensions, which we identified based on the pedagogical piano literature and pilot studies with professional pianists, together present a rich picture of physical gestures at the keyboard, including information on the performer's motion before, during, and after a note is played. User studies confirm that the sensed dimensions are intuitive and controllable and that mappings between gesture and sound produce novel, playable musical instruments, even for users without prior keyboard experience. The multidimensional sensing capability demonstrated in this paper is also potentially applicable to button interfaces outside the musical domain. Andrew P. McPherson, Youngmoo E. Kim |
CHI | 2 |
| 2011 | Robot audition and beat identification in noisy environmentsabstractIn pursuit of our long-term goal of developing an interactive humanoid musician, we are developing robust methods to determine musical beat locations from live acoustic sources. A variety of beat tracking systems have been previously developed, but for the most part they are optimized for direct audio input (no acoustic channel and no noise). The presence of an acoustic channel and noise typically degrades performance substantially. A robot's motors, in particular, create nonstationary noise that can be difficult for a beat detection system to accommodate, Using an algorithm previously developed by the authors, we explore techniques for reducing the effects of the acoustic channel and noise on the system, enabling a humanoid to robustly follow music under realistic conditions. David Grunberg, Daniel M. Lofaro, Paul Y. Oh, Youngmoo E. Kim |
IROS | 4 |
| 2010 | Beat-Sync-Mash-Coder: A web application for real-time creation of beat-synchronous music mashupsabstractWe present the Beat-Sync-Mash-Coder, a new tool for semi-automated real-time creation of beat-synchronous music mashups. We combine phase vocoder and beat tracker technology to automate the task of synchronizing clips. Freeing the user from this task allows us to replace the traditional audio editing paradigm of the Digital Audio Workstation with an intuitive clip selection interface. The application is completely web-based and operates in the ubiquitous cross-platform Flash framework. The efficiency of our implementation is reflected in performance tests, which demonstrate that the system can sustain real-time phase vocoding of 5-9 simultaneous audio signals on consumer-level hardware. This allows the user to easily create dynamic, intricate and musically coherent acoustic soundscapes. Based on an initial user study with 24 high school students, we also find that the Beat-Sync-Mash-Coder is engaging and can get students excited about music and technology. Garth Griffin, Youngmoo E. Kim, Douglas Turnbull |
ICASSP | 2 |
| 2010 | Prediction of Time-Varying Musical Mood Distributions Using Kalman FilteringabstractThe medium of music has evolved specifically for the expression of emotions, and it is natural for us to organize music in terms of its emotional associations. In previous work, we have modeled human response labels to music in the arousal-valence (A-V) representation of affect as a time-varying, stochastic distribution reflecting the ambiguous nature of the perception of mood. These distributions are used to predict A-V responses from acoustic features of the music alone via multi-variate regression. In this paper, we extend our framework to account for multiple regression mappings contingent upon a general location in A-V space. Furthermore, we model A-V state as the latent variable of a linear dynamical system, more explicitly capturing the dynamics of musical mood. We validate this extension using a "genie-bounded" approach, in which we assume that a piece of music is correctly clustered in A-V space a priori, demonstrating significantly higher theoretical performance than the previous single-regressor approach. Erik M. Schmidt, Youngmoo E. Kim |
ICMLA | 2 |
| 2009 | An audio DSP Toolkit for rapid application development in FlashabstractThe Adobe Flash platform has become the de facto standard for developing and deploying media rich Web applications and games. The relative ease-of-development and cross-platform architecture of Flash enables designers to rapidly prototype graphically rich interactive applications, but comprehensive support for audio and signal processing has been lacking. ActionScript, the primary development language used for Flash, is poorly suited for DSP algorithms. To address the inherent challenges in the integration of interactive audio processing into Flash-based applications, we have developed the DSP Audio Toolkit for Flash, which offers significant performance improvements over algorithms implemented in Java or ActionScript. By developing this toolkit, we hope to open up new possibilities for Flash applications and games, enabling them to utilize real-time audio processing as a means to drive gameplay and improve the experience of the end user. Travis M. Doll, Raymond Migneco, Jeffrey J. Scott, Youngmoo E. Kim |
MMSP | 4 |
| 2007 | Blind channel identification for speech dereverberation using l1-norm sparse learningabstractSpeech dereverberation remains an open problem after more than three decades of research. The most challenging step in speech dereverberation is blind chan- nel identification (BCI). Although many BCI approaches have been developed, their performance is still far from satisfactory for practical applications. The main difficulty in BCI lies in finding an appropriate acoustic model, which not only can effectively resolve solution degeneracies due to the lack of knowledge of the source, but also robustly models real acoustic environments. This paper proposes a sparse acoustic room impulse response (RIR) model for BCI, that is, an acous- tic RIR can be modeled by a sparse FIR filter. Under this model, we show how to formulate the BCI of a single-input multiple-output (SIMO) system into a l1- norm regularized least squares (LS) problem, which is convex and can be solved efficiently with guaranteed global convergence. The sparseness of solutions is controlled by l1-norm regularization parameters. We propose a sparse learning scheme that infers the optimal l1-norm regularization parameters directly from microphone observations under a Bayesian framework. Our results show that the proposed approach is effective and robust, and it yields source estimates in real acoustic environments with high fidelity to anechoic chamber measurements. Yuanqing Lin, Jingdong Chen, Youngmoo E. Kim, Daniel D. Lee |
NIPS | 3 |