VLDB 2026 Research / reviewers in the wild / expert
Ping-Keng Jao
dblp:27/9500
· DBLP profile ↗
9ranked-venue papers
7as first author
3since 2021 · last 2023
0000-0003-1715-8472ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorArtificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Audio and music processing · 74% Image and video processing · 26% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › sparse representation
convolutional sparse coding |
0.2 | 1 | 2016 | Monaural Music Source Separation Using Convolutional Sparse Coding · IEEE ACM Trans. Audio Speech Lang. Process. 2016 |
Audio and music processing › source separation
music source separation |
0.2 | 1 | 2016 | Monaural Music Source Separation Using Convolutional Sparse Coding · IEEE ACM Trans. Audio Speech Lang. Process. 2016 |
Audio and music processing
audio representation |
0.2 | 1 | 2014 | AWtoolbox: Characterizing Audio Information Using Audio Words · ACM Multimedia 2014 |
Audio and music processing › music information retrieval
music classification |
0.2 | 1 | 2014 | AWtoolbox: Characterizing Audio Information Using Audio Words · ACM Multimedia 2014 |
Audio and music processing › music transcription
multipitch estimation |
0.1 | 1 | 2016 | Monaural Music Source Separation Using Convolutional Sparse Coding · IEEE ACM Trans. Audio Speech Lang. Process. 2016 |
Methods — techniques the papers use, named apart from their topics
nonnegative matrix factorization · 0.2convolutional sparse coding · 0.2sparse coding · 0.2feature encoding · 0.2dictionary learning · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | EEG-Based Online Regulation of Difficulty in Simulated FlyingabstractAdaptively increasing the difficulty level in learning was shown to be beneficial than increasing the level after some fixed time intervals. To efficiently adapt the level, we aimed at decoding the subjective difficulty level based on Electroencephalography (EEG) signals. We designed a visuomotor learning task that one needed to pilot a simulated drone through a series of waypoints of different sizes, to investigate the effectiveness of the EEG decoder. The EEG decoder was compared with another condition that the subjects decided when to increase the difficulty level. We examined the decoding performance together with behavioral outcomes. The online accuracies were higher than the chance level for 16 out of 26 cases, and the behavioral results, such as task scores, skill curves, and learning patterns, of EEG condition were similar to the condition based on manual regulation of difficulty. Ping-Keng Jao, Ricardo Chavarriaga, José del R. Millán |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Machine-Learning Based Monitoring of Cognitive Workload in Rescue Missions With DronesabstractIn search and rescue missions, drone operations are challenging and cognitively demanding. High levels of cognitive workload can affect rescuers' performance, leading to failure with catastrophic outcomes. To face this problem, we propose a machine learning algorithm for real-time cognitive workload monitoring to understand if a search and rescue operator has to be replaced or if more resources are required. Our multimodal cognitive workload monitoring model combines the information of 25 features extracted from physiological signals, such as respiration, electrocardiogram, photoplethysmogram, and skin temperature, acquired in a noninvasive way. To reduce both subject and day inter-variability of the signals, we explore different feature normalization techniques, and introduce a novel weighted-learning method based on support vector machines suitable for subject-specific optimizations. On an unseen test set acquired from 34 volunteers, our proposed subject-specific model is able to distinguish between low and high cognitive workloads with an average accuracy of 87.3% and 91.2% while controlling a drone simulator using both a traditional controller and a new-generation controller, respectively. Fabio Dell'Agnola, Ping-Keng Jao, Adriana Arza Valdés, Ricardo Chavarriaga, José del R. Millán, Dario Floreano, David Atienza 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | EEG Correlates of Difficulty Levels in Dynamical Transitions of Simulated Flying and Mapping TasksabstractDecoding the subjective perception of task difficulty may help improve operator performance, i.e., automatically optimize the task difficulty level. Here, we aim to decode a compound of cognitive states that covaries with the task difficulty level. We designed a protocol composed of two different subtasks, flying and visual recognition, to induce different difficulty levels. We first showed that electroencephalography (EEG) signals can be a reliable source for discriminating different compound states. To gain insight into the underlying components in the compound states, we examined the attentional index and engagement index as in our previous study. We showed that, first, attention and engagement are essential components but fail to provide the best accuracy, and, second, our model is consistent with our previous study, which means that lateralized modulations in the α bands are representative of the flying task. We also analyzed a practical issue in the design of adaptive human-machine interaction (HMI) systems, namely, the latency of changes in the user's compound state. We hypothesized that the EEG correlates of the task difficulty level do not instantaneously reflect the changes in the task difficulty. We validated the hypothesis by measuring the time required for our decoders to provide stable accuracy after the task changed. This amount of time, or latency, could be as high as ten seconds. The results suggest that the latency of changes in the user's compound state between different tasks is a factor that should be taken into account when building adaptive HMI systems. Ping-Keng Jao, Ricardo Chavarriaga, Fabio Dell'Agnola, Adriana Arza Valdés, David Atienza 0001, José del R. Millán |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2018 | Analysis of EEG Correlates of Perceived Difficulty in Dynamically Changing Flying TasksabstractReal-time workload estimation can be an important tool to improve human-machine interaction. Despite multiple efforts in this sense, most studies focused on tasks designed to have extreme conditions of workload (high vs low), and these levels remain rather constant throughout task execution. However, some applications may induce changing levels of workload and it is not clear how decoders trained in the extreme conditions will perform. In this study we study this scenario in a simulated drone task where participants had to fly through several waypoints. The task difficulty was modulated by the size of waypoints; in one of the studied conditions, the size was modulated in real time by the EEG-based decoding of the perceived difficulty. We show that this protocol can effectively induce different levels of workload and perceived difficulty. Furthermore, post-experiment analysis using a sparse-regularized technique supports the feasibility of decoding perceived difficulty in this dynamically changing task above chance level (average class-balanced accuracy across subjects 67% ± 7%). Interestingly, an analysis of the selected features in independent recording sessions showed that roughly 20% of the features were stable across multiple sessions. Ping-Keng Jao, Ricardo Chavarriaga, José del R. Millán |
SMC | 1 |
| 2016 | Monaural Music Source Separation Using Convolutional Sparse CodingabstractWe present a comprehensive performance study of a new time-domain approach for estimating the components of an observed monaural audio mixture. Unlike existing time-frequency approaches that use the product of a set of spectral templates and their corresponding activation patterns to approximate the spectrogram of the mixture, the proposed approach uses the sum of a set of convolutions of estimated activations with prelearned dictionary filters to approximate the audio mixture directly in the time domain. The approximation problem can be solved by an efficient convolutional sparse coding algorithm. The effectiveness of this approach for source separation of musical audio has been demonstrated in our prior work, but under rather restricted and controlled conditions, requiring the musical score of the mixture being informed a priori and little mismatch between the dictionary filters and the source signals. In this paper, we report an evaluation that considers wider, and more practical, experimental settings. This includes the use of an audio-based multipitch estimation algorithm to replace the musical score, and an external dataset of audio single notes to construct the dictionary filters. Our result shows that the proposed approach remains effective with a larger dictionary, and compares favorably with the state-of-the-art nonnegative matrix factorization approach. However, in the absence of the score and in the case of a small dictionary, our approach may not be better. Ping-Keng Jao, Li Su 0004, Yi-Hsuan Yang, Brendt Wohlberg |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Informed monaural source separation of music based on convolutional sparse codingabstractMonaural source separation is a challenging problem that has many important applications in music information retrieval. In this paper, we focus on the score-informed variant of this problem. While non-negative matrix factorization and some other approaches have been shown effective, few existing approaches have properly taken the phase information into account. There are unnatural sound in the separation result, as the phase of each source signal is considered equivalent to the phase of the mixed signal. To remedy this, we propose to perform source separation directly in the time domain using a convolutional sparse coding (CSC) approach. Evaluation on the Bach10 dataset shows that, when the instrument, pitch and onset/offset time are informed, the source to distortion ratio of the separation result reaches 8.59 dB, which is 2.02 dB higher than a state-of-the-art system called Soundprism. Ping-Keng Jao, Yi-Hsuan Yang, Brendt Wohlberg |
ICASSP | 1 |
| 2015 | Music Annotation and Retrieval using Unlabeled Exemplars: Correlation and Sparse CodesabstractTagging music signals with semantic labels such as genres, moods and instruments is important for content-based music retrieval and recommendation. While considerable effort has been made, automatic music annotation is still considered challenging due to the difficulty of extracting good audio features that capture the characteristics of different tags. To address this issue, we present in this letter two exemplar-based approaches that represent the content of a music clip by referring to a large set of unlabeled audio exemplars. The first approach represents a music clip by the set of audio exemplars that is highly correlated with the short-time feature vectors of the clip, whereas the second approach represents a music clip as sparse linear combinations of its short-time feature vectors over the audio exemplars. Music annotation is then performed by learning the relevance of the audio examples to different tags using labeled data. These two approaches effectively capitalize the availability of unlabeled data to explore the commonality of music signals to find out tag-specific acoustic patterns, without domain knowledge and feature design. Evaluation on the CAL10k music genre tagging dataset for tag-based music retrieval shows that, with thousands of unlabeled audio examples randomly drawn from the Million Song Dataset, the proposed approaches lead to remarkably higher precision rates than existing approaches. Ping-Keng Jao, Yi-Hsuan Yang |
IEEE Signal Process. Lett. | 1 |
| 2014 | Modified lasso screening for audio word-based music classification using large-scale dictionaryabstractRepresenting music information using audio codewords has led to state-of-the-art performance on various music classifcation benchmarks. Comparing to conventional audio descriptors, audio words offer greater fexibility in capturing the nuance of music signals, in that each codeword can be viewed as a quantization of the music universe and that the quantization goes finer as the size of the dictionary (i.e., audio codebook) increases. In practice, however, the high computational cost of codeword assignment might discourage the use of a large dictionary. This paper presents two modifications of a LASSO screening technique developed in the compressive sensing field to speed up the codeword assignment process. The first modification exploits the repetitive nature of music signals, whereas the second one relaxes a screening constraint that is specific to reconstruction but not for classifcation. Our experiments show that the proposed method enables the use of a dictionary of 10,000 codewords with runtime close to the case of using a dictionary of 1,000 codewords. Moreover, using the larger dictionary significantly improves the mean average precision (MAP) from 0.219 to 0.246 for tagging thousands of tracks with 147 possible genre tags. Ping-Keng Jao, Chin-Chia Michael Yeh, Yi-Hsuan Yang |
ICASSP | 1 |
| 2014 | AWtoolbox: Characterizing Audio Information Using Audio WordsabstractThis paper presents the AWtoolbox, an open-source software designed for extracting the audio word (AW) representation of audio signals. The toolbox comes with a graphical user interface that helps a user design custom AW extraction pipelines and various algorithms for feature encoding, dictionary learning, result rectification, pooling, normalization and others. This paper also reports a benchmark comparing eight AW representations computed by the toolbox against state-of-the-art low-level and mid-level timbre, rhythmic and tonal descriptors of music and sound. The evaluation result shows that sparse coding (SC) based AW representation leads to very competitive performances across the three tested sound and music classification tasks. AWtoolbox is available for download at http://mac.citi.sinica.edu.tw/awtoolbox. Chin-Chia Michael Yeh, Ping-Keng Jao, Yi-Hsuan Yang |
ACM Multimedia | 2 |