EDBT 2026 Demo / reviewers in the wild / expert
Yusong Wu
dblp:255/5686
· DBLP profile ↗
10ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FLAM: Frame-Wise Language-Audio ModelingabstractRecent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they still lack fine-grained labeling capability to pinpoint when an event occurs. While traditional sound event detection models can precisely localize events, they are limited to pre-defined categories, making them ineffective for real-world scenarios with out-of-distribution events. In this work, we introduce FLAM, an open-vocabulary contrastive audio-language model capable of localizing specific sound events. FLAM employs a memory-efficient and calibrated frame-wise objective with logit adjustment to address spurious correlations, such as event dependencies and label imbalances during training. To enable frame-wise supervision, we leverage a large-scale dataset with diverse audio events, LLM-generated captions and simulation. Experimental results and case studies demonstrate that FLAM significantly improves the open-vocabulary localization capability while maintaining strong performance in global retrieval and downstream tasks. Yusong Wu, Christos Tsirigotis, Ke Chen 0021, Cheng-Zhi Anna Huang, Aaron C. Courville, Oriol Nieto, Prem Seetharaman, Justin Salamon |
ICML | 1 |
| 2024 | MusicLDM: Enhancing Novelty in text-to-music Generation Using Beat-Synchronous mixup StrategiesabstractDiffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited availability of music data and sensitive issues related to copyright and plagiarism. In this paper, to tackle these challenges, we first construct a state-of-the-art text-to-music model, MusicLDM, that adapts Stable Diffusion and AudioLDM architectures to the music domain. Then, to address the limitations of training data and to avoid plagiarism, we leverage a beat tracking model and propose two different mixup strategies for data augmentation: beat-synchronous audio mixup and beat-synchronous latent mixup, which recombine training audio directly or via a latent embeddings space, respectively. Such mixup strategies encourage the model to interpolate between musical training samples and generate new music within the convex hull of the training data, making the generated music more diverse while still staying faithful to the corresponding style. In addition to popular evaluation metrics, we design several new evaluation metrics based on CLAP score to demonstrate that our proposed MusicLDM and beat-synchronous mixup strategies improve both the quality and novelty of generated music, as well as the correspondence between input text and generated music. Ke Chen 0021, Yusong Wu, Haohe Liu, Marianna Nezhurina, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 2 |
| 2024 | Adaptive Accompaniment with ReaLchordsabstractJamming requires coordination, anticipation, and collaborative creativity between musicians. Current generative models of music produce expressive output but are not able to generate in an online manner, meaning simultaneously with other musicians (human or otherwise). We propose ReaLchords, an online generative model for improvising chord accompaniment to user melody. We start with an online model pretrained by maximum likelihood, and use reinforcement learning to finetune the model for online use. The finetuning objective leverages both a novel reward model that provides feedback on both harmonic and temporal coherency between melody and chord, and a divergence term that implements a novel type of distillation from a teacher model that can see the future melody. Through quantitative experiments and listening tests, we demonstrate that the resulting model adapts well to unfamiliar input and produce fitting accompaniment. ReaLchords opens the door to live jamming, as well as simultaneous co-creation in other modalities. Yusong Wu, Tim Cooijmans, Kyle Kastner, Adam Roberts, Ian Simon, Alexander Scarlatos, Chris Donahue, Cassie Tarakajian, Shayegan Omidshafiei, Aaron C. Courville, Pablo Samuel Castro, Natasha Jaques, Cheng-Zhi Anna Huang |
ICML | 1 |
| 2024 | An 112-Ch Neural Signal Acquisition SoC With Full-Channel Read-Out and Processing AcceleratorsabstractMultichannel neural signal acquisition and processing play a pivotal role in advancing neuroscience research. This article proposes a 112-channel system-on-chip (SoC) design for neural signal acquisition and processing, comprising full-channel read-out circuits, neural signal-processing accelerators, and a 32-bit RISC-V core. A clock-domain-crossing (CDC) structure is devised to minimize data storage overhead in read-out circuits, facilitating comprehensive data acquisition and high-throughput simultaneous transmission from all channels. The channel-specific processing unit incorporates hardware-efficient designs for lossless compression, spike detection, and extraction of spike features. A multistage predictor module serves the dual purpose of narrowing data distribution during compression and signal augmentation during spike detection. The proposed design was fabricated in 40-nm technology with an area of 6.67 mm2. The acquisition of 112 channels achieves a peak data rate of 57.3 Mbps, with a total power consumption of 3.07 mW, wherein 0.87 mW is attributed to the read-out circuits. The processing accelerators feature an area consumption of 0.011 mm2/ch, and a minimal power consumption of only$0.2~{\mu }$W/ch under a 32-kHz clock. The effectiveness of the proposed design is validated through in vivo recording experiments conducted on rats by integrating with flexible implantable electrodes. Zijian Tang, Yongxiang Guo, Minqian Zheng, Yusong Wu, Runjiu Fang, Milin Zhang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption AugmentationabstractContrastive learning has shown remarkable success in the field of multimodal representation learning. In this paper, we propose a pipeline of contrastive language-audio pretraining to develop an audio representation by combining audio data with natural language descriptions. To accomplish this target, we first release LAION-Audio-630K, a large collection of 633,526 audio-text pairs from different data sources. Second, we construct a contrastive language-audio pretraining model by considering different audio encoders and text encoders. We incorporate the feature fusion mechanism and keyword-to-caption augmentation into the model design to further enable the model to process audio inputs of variable lengths and enhance the performance. Third, we perform comprehensive experiments to evaluate our model across three tasks: text-to-audio retrieval, zero-shot audio classification, and supervised audio classification. The results demonstrate that our model achieves superior performance in text-to-audio retrieval task. In audio classification tasks, the model achieves state-of-the-art performance in the zero-shot setting and is able to obtain performance comparable to models’ results in the non-zero-shot setting. LAION-Audio-630K1and the proposed model2are both available to the public. Yusong Wu, Ke Chen 0021, Yuchen Hui, Taylor Berg-Kirkpatrick, Shlomo Dubnov |
ICASSP | 1 |
| 2022 | MIDI-DDSP: Detailed Control of Musical Performance via Hierarchical Modeling
Yusong Wu, Ethan Manilow, Rigel Swavely, Kyle Kastner, Tim Cooijmans, Aaron C. Courville, Cheng-Zhi Anna Huang, Jesse H. Engel |
ICLR | 1 |
| 2022 | A 16-Channel Neural Recorder with 2.8 nJ/bit, 971.4 kbps sub-2.4 GHz polar transmitterabstractThis paper proposed a miniature neural interface system. A single chip neural recording SoC was fabricated in 40nm CMOS process with an area of 3mm×3mm. It integrated a 16-channel analog front end (AFE), and a low power constant envelope polar transmitter. The general form of continuous phase modulation was used as the modulation scheme. Algorithms for receiver including frequency offset calibration, frame synchronization, and symbol demodulation were proposed and implemented on a software-defined radio platform. Simulation results showed that a bit error rate of $10^{-4}$ is achieved at the signal to noise ratio of 19 dB at high data rate mode of 971.4 kbps. A graphic user interface was designed for channel decoding and real-time display. Experimental results showed that the input referred noise of the AFE is 2.87$\mu V_{rms}$, and the energy efficiency of the transmitter is 2. 8nJ/bit. The proposed chip consumes 5. 47mW power in total in its maximum workload. The neural signal can be correctly decoded at least at a RSSI (Received Signal Strength Indicator) of -95dBm, and a working distance of 8 m. In-vivo tests on rat have been conducted, showing a good usability of the proposed system. Heng Huang 0009, Yusong Wu, Xiliang Liu, Zijian Tang, Tianhe Jiang, Xiong Zhong, Milin Zhang 0001 |
ISCAS | 3 |
| 2021 | 3M-AI: A Multi-task and Multi-core Virtualization Framework for Multi-FPGA AI Systems in the CloudabstractWith the ever-growing demands for online Artificial Intelligence (AI), the hardware virtualization support for deep learning accelerators is vital for providing AI capability in the cloud. Three basic features, multi-task, dynamic workload, and remote access, are fundamental for hardware virtualization. However, most of the deep learning accelerators do not support concurrent execution of multiple tasks. Besides, the SOTA multi-DNN scheduling algorithm for NN accelerators neither consider the multi-task concurrent execution and resources allocation for the multi-core DNN accelerators. Moreover, existing GPU virtualized solutions could introduce a huge remote access latency overhead, resulting in a severe system performance drop. Shulin Zeng, Guohao Dai 0001, Hanbo Sun, Jun Liu 0117, Hongren Zheng, Yusong Wu, Yi Cai 0003, Yu Wang 0002, Huazhong Yang |
FPGA | 6 |
| 2020 | Peking Opera Synthesis via Duration Informed Attention NetworkabstractPeking Opera has been the most dominant form of Chinese performing art since around 200 years ago.A Peking Opera singer usually exhibits a very strong personal style via introducing improvisation and expressiveness on stage which leads the actual rhythm and pitch contour to deviate significantly from the original music score.This inconsistency poses a great challenge in Peking Opera singing voice synthesis from a music score.In this work, we propose to deal with this issue and synthesize expressive Peking Opera singing from the music score based on the Duration Informed Attention Network (DurIAN) framework.To tackle the rhythm mismatch, Lagrange multiplier is used to find the optimal output phoneme duration sequence with the constraint of the given note duration from music score.As for the pitch contour mismatch, instead of directly inferring from music score, we adopt a pseudo music score generated from the real singing and feed it as input during training.The experiments demonstrate that with the proposed system we can synthesize Peking Opera singing voice with high-quality timbre, pitch and expressiveness. Yusong Wu, Shengchen Li, Chengzhu Yu, Heng Lu 0004, Chao Weng, Dong Yu 0001 |
INTERSPEECH | 1 |
| 2020 | DurIAN-SC: Duration Informed Attention Network Based Singing Voice Conversion SystemabstractSinging voice conversion is converting the timbre in the source singing to the target speaker's voice while keeping singing content the same. However, singing data for target speaker is much more difficult to collect compared with normal speech this http URL this paper, we introduce a singing voice conversion algorithm that is capable of generating high quality target speaker's singing using only his/her normal speech data. First, we manage to integrate the training and conversion process of speech and singing into one framework by unifying the features used in standard speech synthesis system and singing synthesis system. In this way, normal speech data can also contribute to singing voice conversion training, making the singing voice conversion system more robust especially when the singing database is small.Moreover, in order to achieve one-shot singing voice conversion, a speaker embedding module is developed using both speech and singing data, which provides target speaker identify information during conversion. Experiments indicate proposed sing conversion system can convert source singing to target speaker's high-quality singing with only 20 seconds of target speaker's enrollment speech data. Chengzhu Yu, Heng Lu 0004, Chao Weng, Yusong Wu, Dong Yu 0001 |
INTERSPEECH | 6 |