Jiaqi Su

dblp:243/6787 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Towards Fair and Equitable Incentives to Motivate Paid and Unpaid Crowd Contributions
Shaun Wallace, Talie Massachi, Jiaqi Su, David Bryan Miller, Jeff Huang 0002
CHI3
2025 Unraveling the Effects of Synthetic Data on End-to-End Autonomous Driving
Junhao Ge, Zuhong Liu, Longteng Fan, Jiaqi Su, Zhejun Zhang, Siheng Chen
ICCV5
2024 TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling
abstract
As a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requires not only modeling the relationships between objects in the image but also mining the connections between adjacent images. Recent approaches primarily utilize either end-to-end frameworks or multi-stage frameworks to generate relevant stories, but they usually overlook latent topic information. In this paper, in order to generate a more coherent and relevant story, we propose a novel method, Topic Aware Reinforcement Network for VIsual StoryTelling (TARN-VIST). In particular, we pre-extracted the topic information of stories from both visual and linguistic perspectives. Then we apply two topic-consistent reinforcement learning rewards to identify the discrepancy between the generated story and the human-labeled story so as to refine the whole generation process. Extensive experimental results on the VIST dataset and human evaluation demonstrate that our proposed model outperforms most of the competitive models across multiple evaluation metrics.
Weiran Chen 0001, Jiaqi Su, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu
LREC/COLING3
2024 MDX-GAN: Enhancing Perceptual Quality in Multi-Class Source Separation Via Adversarial Training
abstract
Audio source separation aims to extract individual sound sources from an audio mixture. Recent studies on source separation focus primarily on minimizing signal-level distance, typically measured by source-to-distortion ratio (SDR). However, scant attention has been given to the perceptual quality of the separated tracks. In this paper, we propose MDX-GAN, an efficient and high-fidelity audio source separator based on MDX-Net for multiple sound classes. We leverage different training objectives to enhance the perceptual quality of audio source separation. Specifically, we adopt perceptually-motivated loss functions on top of the waveform loss, including multi-resolution STFT and Mel-spectrogram losses, and employ the adversarial training paradigm with multi-domain and multi-scale discriminators to refine the perceptual quality of separation. Additionally, we extend the model to support multiple sound classes within a single network via feature-wise linear modulation (FiLM). We conduct both objective and subjective experiments to evaluate MDX-GAN on real-world settings, and assess the impacts of design components on the perceptual quality and SDR scores. Results demonstrate that MDX-GAN accurately separates the sound source and achieves superior perceptual quality.
Ke Chen 0021, Jiaqi Su, Zeyu Jin
ICASSP2
2024 Maskmark: Robust Neuralwatermarking for Real and Synthetic Speech
abstract
High-quality speech synthesis models may be used to spread misinformation or impersonate voices. Audio watermarking can combat misuse by embedding a traceable signature in generated audio. However, existing audio watermarks typically demonstrate robustness to only a small set of transformations of the watermarked audio. To address this, we propose MaskMark, a neural network-based digital audio watermarking technique optimized for speech. MaskMark embeds a secret key vector in audio via a multiplicative spectrogram mask, allowing the detection of watermarked speech segments even under substantial signal-processing or neural network-based transformations. Comparisons to a state-of-the-art baseline on natural and synthetic speech corpora and a human subjects evaluation demonstrate MaskMark’s superior robustness in detecting watermarked speech while maintaining high perceptual transparency.
Patrick O'Reilly, Zeyu Jin, Jiaqi Su, Bryan Pardo
ICASSP3
2024 Glocal Cascading Network for Topic Enhanced Visual Storytelling
abstract
As a cross-modal task, visual storytelling aims to generate a semantically coherent story for an ordered image sequence. Despite significant achievements in existing methods for this task, few works focus on improving the conception ability which humans usually use when writing stories. In this work, we propose a framework called GLocal Cascading Network for Topic Enhanced Visual Storytelling which explores the conception ability by pre-modeling a latent topic for each image during story telling. Inspired by the global-local (glocal) ideology, we firstly propose a hierarchical latent-topic decoder consisting of two levels of topic generator which respectively focus on different levels of topic information. Then we propose a topic-aware loss which encourages the model to focus on the topic information of the story. With these two novel modules, our framework can effectively utilize the topic information and improve the informativeness and consistency of stories. Our model has been proven highly competitive across multiple metrics through extensive experiments conducted on the VIST dataset.
Jiaqi Su, Weiran Chen 0001, Yi Ji 0001, Chunping Liu
ICASSP1
2024 GR0: Self-Supervised Global Representation Learning for Zero-Shot Voice Conversion
abstract
Research in generative self-supervised learning (SSL) has largely focused on local embeddings for tokenized sequences. We introduce a generative SSL framework that learns a global representation that is disentangled from local embeddings. We apply this technique to jointly learn a global speaker embedding and a zero-shot voice converter. The converter modifies recorded speech to sound as if it were spoken by a different person while preserving the content, using only a short reference clip unavailable to the model during training. Listening experiments conducted on an unseen dataset show that our models significantly outperform SOTA baselines in both quality and speaker similarity for various datasets and unseen languages.
Jiaqi Su, Adam Finkelstein, Zeyu Jin
ICASSP2
2024 Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation
Ke Chen 0021, Jiaqi Su, Taylor Berg-Kirkpatrick, Shlomo Dubnov, Zeyu Jin
INTERSPEECH2
2024 Genhancer: High-Fidelity Speech Enhancement via Generative Modeling on Discrete Codec Tokens
Haici Yang, Jiaqi Su, Zeyu Jin
INTERSPEECH2
2024 DeepSS2GO: protein function prediction from secondary structure
abstract
Predicting protein function is crucial for understanding biological life processes, preventing diseases and developing new drug targets. In recent years, methods based on sequence, structure and biological networks for protein function annotation have been extensively researched. Although obtaining a protein in three-dimensional structure through experimental or computational methods enhances the accuracy of function prediction, the sheer volume of proteins sequenced by high-throughput technologies presents a significant challenge. To address this issue, we introduce a deep neural network model DeepSS2GO (Secondary Structure to Gene Ontology). It is a predictor incorporating secondary structure features along with primary sequence and homology information. The algorithm expertly combines the speed of sequence-based information with the accuracy of structure-based features while streamlining the redundant data in primary sequences and bypassing the time-consuming challenges of tertiary structure analysis. The results show that the prediction performance surpasses state-of-the-art algorithms. It has the ability to predict key functions by effectively utilizing secondary structure information, rather than broadly predicting general Gene Ontology terms. Additionally, DeepSS2GO predicts five times faster than advanced algorithms, making it highly applicable to massive sequencing data. The source code and trained models are available at https://github.com/orca233/DeepSS2GO.
Fu V. Song, Jiaqi Su, Sixing Huang, Kaiyue Li, Maofu Liao
Briefings Bioinform.2
2024 Quality evaluation methods of handwritten Chinese characters: a comprehensive survey
Weiran Chen 0001, Jiaqi Su, Guiqian Zhu, Ying Li 0065, Yi Ji 0001, Chunping Liu
Multim. Syst.2
2022 Controllable Speech Representation Learning Via Voice Conversion and AIC Loss
abstract
Speech representation learning transforms speech into features that are suitable for downstream tasks, e.g. speech recognition, phoneme classification, or speaker identification. For such recognition tasks, a representation can be lossy (non-invertible), which is typical of BERT-like self-supervised models. However, when used for synthesis tasks, we find these lossy representations prove to be insufficient to plausibly reconstruct the input signal. This paper introduces a method for invertible and controllable speech representation learning based on disentanglement. The representation can be decoded into a signal perceptually identical to the original. Moreover, its disentangled components (content, pitch, speaker identity, and energy) can be controlled independently to alter the synthesis result. Our model builds upon a zero-shot voice conversion model AutoVC-F0, in which we introduce alteration invariant content loss (AIC loss) and adversarial training (GAN). Through objective measures and subjective tests, we show that our formulation offers significant improvement in voice conversion sound quality as well as more precise control over the disentangled features.
Jiaqi Su, Adam Finkelstein, Zeyu Jin
ICASSP2
2021 Bandwidth Extension is All You Need
abstract
Speech generation and enhancement have seen recent breakthroughs in quality thanks to deep learning. These methods typically operate at a limited sampling rate of 16-22kHz due to computational complexity and available datasets. This limitation imposes a gap between the output of such methods and that of high-fidelity (≥44kHz) real-world audio applications. This paper proposes a new bandwidth extension (BWE) method that expands 8-16kHz speech signals to 48kHz. The method is based on a feed-forward WaveNet architecture trained with a GAN-based deep feature loss. A mean-opinion-score (MOS) experiment shows significant improvement in quality over state-of-the-art BWE methods. An AB test reveals that our 16-to-48kHz BWE is able to achieve fidelity that is typically indistinguishable from real high-fidelity recordings. We use our method to enhance the output of recent speech generation and denoising methods, and experiments demonstrate significant improvement in sound quality over these baselines. We propose this as a general approach to narrow the gap between generated speech and recorded speech, without the need to adapt such methods to higher sampling rates.
Jiaqi Su, Adam Finkelstein, Zeyu Jin
ICASSP1
2020 Acoustic Matching By Embedding Impulse Responses
abstract
The goal of acoustic matching is to transform an audio recording made in one acoustic environment to sound as if it had been recorded in a different environment, based on reference audio from the target environment. This paper introduces a deep learning solution for two parts of the acoustic matching problem. First, we characterize acoustic environments by mapping audio into a low-dimensional embedding invariant to speech content and speaker identity. Next, a waveform-to-waveform neural network conditioned on this embedding learns to transform an input waveform to match the acoustic qualities encoded in the target embedding. Listening tests on both simulated and real environments show that the proposed approach improves on state-of-the-art baseline methods.
Jiaqi Su, Zeyu Jin, Adam Finkelstein
ICASSP1
2020 HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
abstract
Real-world audio recordings are often degraded by factors such as noise, reverberation, and equalization distortion.This paper introduces HiFi-GAN, a deep learning method to transform recorded speech to sound as though it had been recorded in a studio.We use an end-to-end feed-forward WaveNet architecture, trained with multi-scale adversarial discriminators in both the time domain and the time-frequency domain.It relies on the deep feature matching losses of the discriminators to improve the perceptual quality of enhanced speech.The proposed model generalizes well to new speakers, new speech content, and new environments.It significantly outperforms state-of-the-art baseline methods in both objective and subjective experiments.
Jiaqi Su, Zeyu Jin, Adam Finkelstein
INTERSPEECH1
2019 Learning Bandwidth Expansion Using Perceptually-motivated Loss
abstract
We introduce a perceptually motivated approach to bandwidth expansion for speech. Our method pairs a new 3-way split variant of the FFTNet neural vocoder structure with a perceptual loss function, combining objectives from both the time and frequency domains. Mean opinion score tests show that it outperforms baseline methods from both domains, even for extreme bandwidth expansion.
Berthy Feng, Zeyu Jin, Jiaqi Su, Adam Finkelstein
ICASSP3
2019 Perceptually-motivated Environment-specific Speech Enhancement
abstract
This paper introduces a deep learning approach to enhance speech recordings made in a specific environment. A single neural network learns to ameliorate several types of recording artifacts, including noise, reverberation, and non-linear equalization. The method relies on a new perceptual loss function that combines adversarial loss with spectrogram features. Both subjective and objective evaluations show that the proposed approach improves on state-of-the-art baseline methods.
Jiaqi Su, Adam Finkelstein, Zeyu Jin
ICASSP1