Jiangchuan Li

dblp:35/11184 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image Compression
abstract
Neural image compression often faces a challenging trade-off among rate, distortion and perception. While most existing methods typically focus on either achieving high pixel-level fidelity or optimizing for perceptual metrics, we propose a novel approach that simultaneously addresses both aspects for a fixed neural image codec. Specifically, we introduce a plug-and-play module at the decoder side that leverages a latent diffusion process to transform the decoded features, enhancing either low distortion or high perceptual quality without altering the original image compression codec. Our approach facilitates fusion of original and transformed features without additional training, enabling users to flexibly adjust the balance between distortion and perception during inference. Extensive experimental results demonstrate that our method significantly enhances the pretrained codecs with a wide, adjustable distortion-perception range while maintaining their original compression capabilities. For instance, we can achieve more than 150% improvement in LPIPS-BDRate without sacrificing more than 1 dB in PSNR.
Chuqin Zhou, Guo Lu, Jiangchuan Li, Zhengxue Cheng, Li Song 0001, Wenjun Zhang 0001
AAAI3
2025 Differentiable VMAF: A trainable metric for optimizing video compression codec
abstract
Video Multi-method Assessment Fusion (VMAF) is a widely used objective evaluation metric that has shown a stronger correlation with human visual system than other metrics. Some studies have already integrated VMAF as a measure of perceptual quality in the rate-distortion optimization of hybrid codecs. However, VMAF has long been considered non-differentiable, limiting its application in learning-based video compression. In this work, we successfully achieve a differentiable VMAF through mathematical approximations. Our implementation of differentiable VMAF maintains the performance of the original VMAF while ensuring accurate gradient calculations. By integrating it into the loss function for optimizing learning-based video compression, our experimental results demonstrate that this approach significantly enhances VMAF performance and improves perceptual quality at all bitrates.
Jiangchuan Li, Chuqin Zhou, Yunuo Chen 0002, Guo Lu
ISCAS1
2025 Instance-Adaptive Spatial-Temporal Enhancement for Efficient Video Compression
abstract
Efficiently compressing HD/UHD content has long been challenging due to high bitrate costs. Instance-adaptive enhancement methods try to tackle this issue by compressing a video at reduced resolution and enhancing it using a neural model specifically overfitted for this video. However, existing methods focus solely on spatial super-resolution (SR) and under-utilize the videos' temporal redundancy. Their limited management of the model's updated parameters also causes excessive overfitting overheads. Therefore, this paper introduces IASTE, the first instance-adaptive enhancement method based on spatial-temporal enhancement (STE), and incorporates low-rank adaptation (LoRA) for efficient model overfitting. Specifically, we downscale videos spatially and temporally to reduce the data volume and achieve efficient video compression. Then, we overfit a specific STE model for each video and use it to enhance the decoded video's spatiotemporal resolution. Leveraging the video swin transformer's strong capability in capturing spatiotemporal correlations, we design a lightweight and efficient model to implement video STE. The model is overfitted for each video using LoRA. By freezing the pre-trained model and selectively updating a few low-rank matrices, the bitrate overhead for model storage can be mitigated. Experiments prove that compared to directly compressing high-frame-rate (HFR), high-resolution (HR) videos, our method achieves around 30% BD-Rate gains on the CTC and UVG datasets, about 15% gains on the YoutubeUGC dataset, and about 10% gains on the ultra-long videos in the Xiph dataset.
Yan Zhao 0041, Zhengxue Cheng, Jiangchuan Li, Donghui Feng 0003, Qunshan Gu, Cheems Wang, Guo Lu, Li Song 0001
IEEE Trans. Image Process.3
2023 Content-adaptive Adversarial Embedding for Image Steganography Using Deep Reinforcement Learning
abstract
Recently, adversarial perturbations have been used to reassign cost which can enhance the security of steganography, called as adversarial embedding. However, existing methods selected costs to be modified by self-defined rules which were hard to achieve the optimal security against steganalyzers. In this paper, we propose an automatic adversarial embedding scheme called RLAE (deep Reinforcement Learning-based content-adaptive Adversarial Embedding). In RLAE, an agent network utilizes a generative network which generates an embedding policy for cost reassignment automatically according to a basic steganography cost map. Then, an environment network employs a steganalyzer as an attack target that offers rewards for optimizing the agent network. To provide more comprehensive information, we design a joint reward by considering both the adversarial perturbations calculated from the environment network and noise residual signal representing image textures. Experimental results show that the security of the proposed RLAE is superior than state-of-the-art works, especially steganography with for the large payloads.
Jie Luo 0005, Peisong He, Hongxia Wang 0001, Chunwang Wu, Wanjie Li, Jiangchuan Li
ICME8
2022 Hierarchical Prosody Modeling and Control in Non-Autoregressive Parallel Neural TTS
abstract
Neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the synthetic speech often represents the average prosodic style of the database instead of having more versatile prosodic variation. Moreover, many models lack the ability to control the output prosody, which does not allow for different styles for the same text input. In this work, we train a non-autoregressive parallel neural TTS front-end model hierarchically conditioned on both coarse and fine-grained acoustic speech features to learn a latent prosody space with intuitive and meaningful dimensions. Experiments show that a non-autoregressive TTS model hierarchically conditioned on utterance-wise pitch, pitch range, duration, energy, and spectral tilt can effectively control each prosodic dimension, generate a wide variety of speaking styles, and provide word-wise emphasis control, while maintaining equal or better quality to the baseline model.
Tuomo Raitio, Jiangchuan Li, Shreyas Seshadri
ICASSP2
2022 Vocal effort modeling in neural TTS for improving the intelligibility of synthetic speech in noise
abstract
We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise.The method consists of first measuring the spectral tilt of unlabeled conventional speech data, and then conditioning a neural TTS model with normalized spectral tilt among other prosodic factors.Changing the spectral tilt parameter and keeping other prosodic factors unchanged enables effective vocal effort control at synthesis time independent of other prosodic factors.By extrapolation of the spectral tilt values beyond what has been seen in the original data, we can generate speech with high vocal effort levels, thus improving the intelligibility of speech in the presence of masking noise.We evaluate the intelligibility and quality of normal speech and speech with increased vocal effort in the presence of various masking noise conditions, and compare these to well-known speech intelligibility-enhancing algorithms.The evaluations show that the proposed method can improve the intelligibility of synthetic speech with little loss in speech quality.
Tuomo Raitio, Petko Petkov, Jiangchuan Li, P. V. Muhammed Shifas, Andrea Davis, Yannis Stylianou
INTERSPEECH3
2022 Emphasis Control for Parallel Neural TTS
abstract
Recent parallel neural text-to-speech (TTS) synthesis methods are able to generate speech with high fidelity while maintaining high performance.However, these systems often lack control over the output prosody, thus restricting the semantic information conveyable for a given text.This paper proposes a hierarchical parallel neural TTS system for prosodic emphasis control by learning a latent space that directly corresponds to a change in emphasis.Three candidate features for the latent space are compared: 1) Variance of pitch and duration within words in a sentence, 2) Wavelet-based feature computed from pitch, energy, and duration, and 3) Learned combination of the two aforementioned approaches.At inference time, word-level prosodic emphasis is achieved by increasing the feature values of the latent space for the given words.Experiments show that all the proposed methods are able to achieve the perception of increased emphasis with little loss in overall quality.Moreover, emphasized utterances were preferred in a pairwise comparison test over the non-emphasized utterances, indicating promise for real-world applications.
Shreyas Seshadri, Tuomo Raitio, Dan Castellani, Jiangchuan Li
INTERSPEECH4
2021 On-Device Neural Speech Synthesis
abstract
Recent advances in text-to-speech (TTS) synthesis, such as Tacotron and WaveRNN, have made it possible to construct a fully neural network based TTS system, by coupling the two components together. Such a system is conceptually sim-ple as it only takes grapheme or phoneme input, uses Mel-spectrogram as an intermediate feature, and directly generates speech samples. The system achieves quality equal or close to natural speech. However, the high computational cost of the system and issues with robustness have limited their usage in real-world speech synthesis applications and products. In this paper, we present key modeling improvements and optimization strategies that enable deploying these models, not only on GPU servers, but also on mobile devices. The proposed system can generate high-quality 24 kHz speech at 5x faster than real time on server and 3x faster than real time on mobile devices.
Sivanand Achanta, Albert Antony, Ladan Golipour, Jiangchuan Li, Tuomo Raitio, Ramya Rasipuram, Jennifer Shi, Jaimin Upadhyay, David Winarsky, Hepeng Zhang
ASRU4
2017 Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System
Tim Capes, Paul Coles, Alistair Conkie, Ladan Golipour, Abie Hadjitarkhani, Qiong Hu 0003, Nancy Huddleston, Melvyn Hunt, Jiangchuan Li, Matthias Neeracher, Kishore Prahallad, Tuomo Raitio, Ramya Rasipuram, Greg Townsend, Becci Williamson, David Winarsky, Zhizheng Wu 0001, Hepeng Zhang
INTERSPEECH9