Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wei Tsung Lu

dblp:230/1613 · also Wei-Tsung Lu · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Audio and music processing · 100%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › music information retrieval › music alignment
lyrics-to-audio alignment
0.712023
ALCAP: Alignment-Augmented Music Captioner · EMNLP 2023
Audio and music processing › music information retrieval
music captioning
0.712023
ALCAP: Alignment-Augmented Music Captioner · EMNLP 2023
Audio and music processing › music information retrieval
music understanding
0.712023
ALCAP: Alignment-Augmented Music Captioner · EMNLP 2023
Audio and music processing › music generation
music style transfer
0.512021
Actions Speak Louder than Listening: Evaluating Music Style Transfer based on Editing Experience · ACM Multimedia 2021
Usability and user experience research
user experience evaluation
0.112021
Actions Speak Louder than Listening: Evaluating Music Style Transfer based on Editing Experience · ACM Multimedia 2021

Methods — techniques the papers use, named apart from their topics

user study · 1.0multimodal alignment · 0.7contrastive learning · 0.7
YearPublicationVenuePosition
2024 Music Source Separation With Band-Split Rope Transformer
abstract
Music source separation (MSS) aims to separate a music recording into multiple musically distinct stems, such as vocals, bass, drums, and more. Recently, deep learning approaches such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have been used, but the improvement is still limited. In this paper, we propose a novel frequency-domain approach (called BS-RoFormer) based on a Band-Split RoPE Transformer architecture. BS-RoFormer relies on a band-split module to project the input complex spectrogram into subband-level representations, and then arranges a stack of hierarchical Transformers to model the inner-band as well as inter-band sequences for multi-band mask estimation. To facilitate training the model for MSS, we propose to use the Rotary Position Embedding (RoPE). The BS-RoFormer system trained on MUSDB18HQ and 500 extra songs ranked the first place in the Music Separation contest of Sound Demixing Challenge (SDX’23). Benchmarking a smaller version of BS-RoFormer on MUSDB18HQ, we achieve state-of-the-art result without extra training data, with 9.80 dB of average SDR.
Wei Tsung Lu, Ju-Chiang Wang, Qiuqiang Kong, Yun-Ning Hung
ICASSP1
2023 ALCAP: Alignment-Augmented Music Captioner
abstract
Music captioning has gained significant attention in the wake of the rising prominence of streaming media platforms.Traditional approaches often prioritize either the audio or lyrics aspect of the music, inadvertently ignoring the intricate interplay between the two.However, a comprehensive understanding of music necessitates the integration of both these elements.In this study, we delve into this overlooked realm by introducing a method to systematically learn multimodal alignment between audio and lyrics through contrastive learning.This not only recognizes and emphasizes the synergy between audio and lyrics but also paves the way for models to achieve deeper cross-modal coherence, thereby producing high-quality captions.We provide both theoretical and empirical results demonstrating the advantage of the proposed method, which achieves new state-of-the-art on two music captioning datasets.
Weituo Hao, Wei Tsung Lu, Changyou Chen, Kristina Lerman, Xuchen Song
EMNLP3
2023 Multitrack Music Transcription with a Time-Frequency Perceiver
abstract
Multitrack music transcription aims to transcribe a music audio input into the musical notes of multiple instruments simultaneously. It is a very challenging task that typically requires a more complex model to achieve satisfactory result. In addition, prior works mostly focus on transcriptions of regular instruments, however, neglecting vocals, which are usually the most important signal source if present in a piece of music. In this paper, we propose a novel deep neural network architecture, Perceiver TF, to model the time-frequency representation of audio input for multitrack transcription. Perceiver TF augments the Perceiver architecture by introducing a hierarchical expansion with an additional Transformer layer to model temporal coherence. Accordingly, our model inherits the benefits of Perceiver that posses better scalability, allowing it to well handle transcriptions of many instruments in a single model. In experiments, we train a Perceiver TF to model 12 instrument classes as well as vocal in a multi-task learning manner. Our result demonstrates that the proposed system outperforms the state-of-the-art counterparts (e.g., MT3 and SpecTNT) on various public datasets.
Wei Tsung Lu, Ju-Chiang Wang, Yun-Ning Hung
ICASSP1
2022 Modeling Beats and Downbeats with a Time-Frequency Transformer
abstract
Transformer is a successful deep neural network (DNN) architecture that has shown its versatility not only in natural language processing but also in music information retrieval (MIR). In this paper, we present a novel Transformer-based approach to tackle beat and downbeat tracking. This approach employs SpecTNT (Spectral- Temporal Transformer in Transformer), a variant of Transformer that models both spectral and temporal dimensions of a time-frequency input of music audio. A SpecTNT model uses a stack of blocks, where each consists of two levels of Transformer encoders. The lower-level (or spectral) encoder handles the spectral features and enables the model to pay attention to harmonic components of each frame. Since downbeats indicate bar boundaries and are often accompanied by harmonic changes, this step may help downbeat modeling. The upper-level (or temporal) encoder aggregates useful local spectral information to pay attention to beat/downbeat positions. We also propose an architecture that combines SpecTNT with a state-of- the-art model, Temporal Convolutional Networks (TCN), to further improve the performance. Extensive experiments demonstrate that our approach can significantly outperform TCN in downbeat tracking while maintaining comparable result in beat tracking.
Yun-Ning Hung, Ju-Chiang Wang, Xuchen Song, Wei Tsung Lu, Minz Won
ICASSP4
2021 Actions Speak Louder than Listening: Evaluating Music Style Transfer based on Editing Experience
Wei Tsung Lu, Meng-Hsuan Wu, Yuh-Ming Chiu, Li Su 0004
ACM Multimedia1