Alexander Lerch 0001

dblp:34/1665 · also Alex Lerch 0001 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0001-6319-578XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
YearPublicationVenuePosition
2025 Uncertainty Estimation in the Real World: A Study on Music Emotion Recognition
Karn Watcharasupat, Yiwei Ding, T. Aleksandra Ma, Pavan Seshadri, Alexander Lerch 0001
ECIR (2)5
2024 ASPED: An Audio Dataset for Detecting Pedestrians
abstract
We introduce the new audio analysis task of pedestrian detection and present a new large-scale dataset for this task. While the preliminary results prove the viability of using audio approaches for pedestrian detection, they also show that this challenging task cannot be easily solved with standard approaches.
Pavan Seshadri, Chaeyeon Han, Bon-Woo Koo, Noah Posner, Subhrajit Guhathakurta, Alexander Lerch 0001
ICASSP6
2024 Quantifying Spatial Audio Quality Impairment
abstract
Spatial audio quality is a highly multifaceted concept, with many interactions between environmental, geometrical, anatomical, psychological, and contextual factors. Methods for characterization or evaluation of the geometrical components of spatial audio quality, however, remain scarce, despite being perhaps the least subjective aspect of spatial audio quality to quantify. By considering interchannel time and level differences relative to a reference signal, it is possible to construct a signal model to isolate some of the spatial distortion. By using a combination of least-square optimization and heuristics, we propose a signal decomposition method to isolate the spatial error, in terms of interchannel gain leakages and changes in relative delays, from a processed signal. This allows the computation of simple energy-ratio metrics, providing objective measures of spatial and non-spatial signal qualities, with minimal assumptions and no dataset dependency. Experiments demonstrate the robustness of the method against common spatial signal degradation introduced by, e.g., audio compression and music source separation.
Karn Watcharasupat, Alexander Lerch 0001
ICASSP2
2024 A Large-Scale Multiobjective Particle Swarm Optimizer With Enhanced Balance of Convergence and Diversity
abstract
Large-scale multiobjective optimization problems (LSMOPs) continue to be challenging for existing multiobjective evolutionary algorithms (MOEAs). The main difficulties are that: 1) the diversity preservation in both the objective space and the decision space needs to be taken into account when solving LSMOPs and 2) the existing learning structures in current MOEAs usually make the learning operators only coincidentally serve convergence and diversity, leading to difficulties in balancing these two factors. Therefore, balancing convergence and diversity in current MOEAs is difficult. To address these issues, this article proposes a multiobjective particle swarm optimizer with enhanced balance of convergence and diversity (MPSO-EBCD). In MPSO-EBCD, a novel velocity update structure for multiobjective particle swarm optimization is put forward, dividing the convergence, and diversity preservation operations into independent components. Following the proposed update structure, a weighted convergence factor is introduced to serve the convergence strategy, whilst a diversity preservation strategy is built to uniformly distribute the particles in the searched space based on a proposed multidimensional local sparseness degree indicator. By this means, MPSO-EBCD is able to balance convergence and diversity with specific parameters in independent operators. Experimental results on LSMOP benchmarks and a voltage transformer optimization problem demonstrate the competitiveness of the proposed algorithm compared to several state-of-the-art MOEAs.
Lei Wang 0006, Li Li 0008, Weian Guo, Qidi Wu, Alexander Lerch 0001
IEEE Trans. Cybern.6
2023 Low-Resource Music Genre Classification with Cross-Modal Neural Model Reprogramming
abstract
Transfer learning (TL) approaches have shown promising results when handling tasks with limited training data. However, considerable memory and computational resources are often required for fine-tuning pre-trained neural networks with target domain data. In this work, we introduce a novel method for leveraging pre-trained speech models for low-resource music classification based on the concept of Neural Model Reprogramming (NMR). NMR aims at re-purposing a pre-trained model from a source domain to a target domain by modifying the input of a frozen pre-trained models for cross-modal adaptation. In addition to the known, input-independent, re-programming method, we propose an new reprogramming paradigm: Input-dependent NMR, to increase adaptability to complex input data such as musical audio. Experimental results suggest that a neural model pre-trained on large-scale datasets can successfully perform music genre classification by using this reprogramming method. The two proposed Input-dependent NMR TL methods outperform fine-tuning-based TL methods on a small genre classification dataset.
Yun-Ning Hung, Chao-Han Huck Yang, Alexander Lerch 0001
ICASSP4
2023 Music Instrument Classification Reprogrammed
Hsin-Hung Chen, Alexander Lerch 0001
MMM (1)2
2022 Deep reinforcement learning for urban multi-taxis cruising strategy
Weian Guo, Zhenyao Hua, Zecheng Kang, Lei Wang 0006, Qidi Wu, Alexander Lerch 0001
Neural Comput. Appl.7
2021 Mind the Beat: Detecting Audio Onsets from EEG Recordings of Music Listening
abstract
We propose a deep learning approach to predicting audio event onsets in electroencephalogram (EEG) recorded from users as they listen to music. We use a publicly available dataset containing ten contemporary songs and concurrently recorded EEG. We generate a sequence of onset labels for the songs in our dataset and trained neural networks (a fully connected network (FCN) and a recurrent neural network (RNN)) to parse one second windows of input EEG to predict one second windows of onsets in the audio. We compare our RNN network to both the standard spectral-flux based novelty function and the FCN. We find that our RNN was able to produce results that reflected its ability to generalize better than the other methods.Since there are no pre-existing works on this topic, the numbers presented in this paper may serve as useful benchmarks for future approaches to this research problem.
Ashvala Vinay, Alexander Lerch 0001, Grace Leslie
ICASSP2
2021 Semi-Supervised Audio Classification with Partially Labeled Data
abstract
Audio classification has seen great progress with the increasing availability of large-scale datasets. These large datasets, however, are often only partially labeled as collecting full annotations is a tedious and expensive process. This paper presents two semi-supervised methods capable of learning with missing labels and evaluates them on two publicly available, partially labeled datasets. The first method relies on label enhancement by a two-stage teacher-student learning process, while the second method utilizes the mean teacher semi-supervised learning algorithm. Our results demonstrate the impact of improperly handling missing labels and compare the benefits of using different strategies leveraging data with few labels. Methods capable of learning with partially labeled data have the potential to improve models for audio classification by utilizing even larger amounts of data without the need for complete annotations.
Siddharth Gururani, Alexander Lerch 0001
ISM2
2021 Attribute-based regularization of latent spaces for variational auto-encoders
Kumar Ashis Pati, Alexander Lerch 0001
Neural Comput. Appl.2
2020 Melody-Conditioned Lyrics Generation with SeqGANs
abstract
Automatic lyrics generation has received attention from both music and AI communities for years. Early rule-based approaches have -due to increases in computational power and evolution in data-driven models mostly been replaced with deep-learning-based systems. Many existing approaches, however, either rely heavily on prior knowledge in music and lyrics writing or oversimplify the task by largely discarding melodic information and its relationship with the text. We propose an end-to-end melody-conditioned lyrics generation system based on Sequence Generative Adversarial Networks (SeqGAN), which generates a line of lyrics given the corresponding melody as the input. Furthermore, we investigate the performance of the generator with an additional input condition: the theme or overarching topic of the lyrics to be generated. We show that the input conditions have no negative impact on the evaluation metrics while enabling the network to produce more meaningful results.
Alexander Lerch 0001
ISM2
2020 Remixing Music with Visual Conditioning
abstract
We propose a visually conditioned music remixing system by incorporating deep visual and audio models. The method is based on a state of the art audio-visual source separation model which performs music instrument source separation with video information. We modified the model to work with user-selected images instead of videos as visual input during inference to enable separation of audio-only content. Furthermore, we propose a remixing engine that generalizes the task of source separation into music remixing. The proposed method is able to achieve improved audio quality compared to remixing performed by the separate-and-add method with a state-of-the-art audiovisual source separation model.
Li-Chia Yang, Alexander Lerch 0001
ISM2
2020 On the evaluation of generative models in music
Li-Chia Yang, Alexander Lerch 0001
Neural Comput. Appl.2
2019 Tuning Frequency Dependency in Music Classification
abstract
Deep architectures have become ubiquitous in Music Information Retrieval (MIR) tasks, however, concurrent studies still lack a deep understanding of the input properties being evaluated by the networks. In this study, we show by the example of a Music Genre Classification system the potential dependency on the tuning frequency, an irrelevant and confounding variable. We generate adversarial samples through pitch-shifting the audio data and investigate the classification accuracy of the output depending on the pitch shift. We find the accuracy to be periodic with a period of one semitone, indicating that the system is utilizing tuning information. We show that proper data augmentation including pitch-shifts smaller than one semitone helps minimizing this problem and point out the need for carefully designed augmentation procedures in related MIR tasks.
Alexander Lerch 0001
ICASSP2
2018 A Review of Automatic Drum Transcription
abstract
In Western popular music, drums and percussion are an important means to emphasize and shape the rhythm, often defining the musical style. If computers were able to analyze the drum part in recorded music, it would enable a variety of rhythm-related music processing tasks. Especially the detection and classification of drum sound events by computational methods is considered to be an important and challenging research problem in the broader field of music information retrieval. Over the last two decades, several authors have attempted to tackle this problem under the umbrella term automatic drum transcription (ADT). This paper presents a comprehensive review of ADT research, including a thorough discussion of the task-specific challenges, categorization of existing techniques, and evaluation of several state-of-the-art systems. To provide more insights on the practice of ADT systems, we focus on two families of ADT techniques, namely methods based on non-negative matrix factorization and recurrent neural networks. We explain the methods’ technical details and drum-specific variations and evaluate these approaches on publicly available data sets with a consistent experimental setup. Finally, the open issues and underexplored areas in ADT research are identified and discussed, providing future directions in this field.
Chih-Wei Wu, Christian Dittmar, Carl Southall, Richard Vogl, Gerhard Widmer, Jason Hockman, Meinard Müller, Alexander Lerch 0001
IEEE ACM Trans. Audio Speech Lang. Process.8
2017 Learning to Fuse Music Genres with Generative Adversarial Dual Learning
abstract
FusionGAN is a novel genre fusion framework for music generation that integrates the strengths of generative adversarial networks and dual learning. In particular, the proposed method offers a dual learning extension that can effectively integrate the styles of the given domains. To efficiently quantify the difference among diverse domains and avoid the vanishing gradient issue, FusionGAN provides a Wasserstein based metric to approximate the distance between the target domain and the existing domains. Adopting the Wasserstein distance, a new domain is created by combining the patterns of the existing domains using adversarial learning. Experimental results on public music datasets demonstrated that our approach could effectively merge two genres.
Zhiqian Chen, Chih-Wei Wu, Yen-Cheng Lu, Alexander Lerch 0001, Chang-Tien Lu
ICDM4
2016 An Unsupervised Approach to Anomaly Detection in Music Datasets
abstract
This paper presents an unsupervised method for systematically identifying anomalies in music datasets. The model integrates categorical regression and robust estimation techniques to infer anomalous scores in music clips. When applied to a music genre recognition dataset, the new method is able to detect corrupted, distorted, or mislabeled audio samples based on commonly used features in music information retrieval. The evaluation results show that the algorithm outperforms other anomaly detection methods and is capable of finding problematic samples identified by human experts. The proposed method introduces a preliminary framework for anomaly detection in music data that can serve as a useful tool to improve data integrity in the future.
Yen-Cheng Lu, Chih-Wei Wu, Chang-Tien Lu, Alexander Lerch 0001
SIGIR4
2011 Strategies for orca call retrieval to support collaborative annotation of a large archive
abstract
The Orchive is a large audio archive of hydrophone recordings of Killer whale (Orcinus orca) vocalizations. Researchers and users from around the world can interact with the archive using a collaborative web-based annotation, visualization and retrieval interface. In addition a mobile client has been written in order to crowdsource Orca call annotation. In this paper we describe and compare different strategies for the retrieval of discrete Orca calls. In addition, the results of the automatic analysis are integrated in the user interface facilitating annotation as well as leveraging the existing annotations for supervised learning. The best strategy achieves a mean average precision of 0.77 with the first retrieved item being relevant 95% of the time in a dataset of 185 calls belonging to 4 types.
Steven R. Ness, Alexander Lerch 0001, George Tzanetakis
MMSP2