VLDB 2026 Research / reviewers in the wild / expert
Yicheng Gu
dblp:359/2013
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Yiming Zhang 0031, Yicheng Gu, Yanhong Zeng, Zhening Xing, Yuancheng Wang, Zhizheng Wu 0001, Kai Chen 0026 |
Int. J. Comput. Vis. | 2 |
| 2025 | Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN
Yicheng Gu, Chaoren Wang, Zhizheng Wu 0001, Lauri Juvela |
INTERSPEECH | 1 |
| 2025 | gFlow: Distributed Real-Time Reverse Remote Rendering System Model
Yixiao Xu, Wanzhao Xu, Yicheng Gu, Yun Wang 0039, Jiangyuan Ma, Zhengwei Qi |
MMM (2) | 4 |
| 2025 | gCom: Fine-grained Compressors in Graphics Memory of Mobile GPUabstractToday, GPUs significantly boost rendering performance. However, the high memory requirements limit their use, especially on low-end mobile platforms. Compression techniques have been widely adopted to reduce memory consumption but face two primary issues when applied to mobile GPUs: (1) low repetition ratio caused by small raw data sizes and concurrency, and (2) low locality caused by unpredictable rendering behaviors. These two limitations result in a low compression ratio when compressors are applied to low-end mobile devices. This article introduces gCom , a fine-grained rendering compressor accelerated by GPUs. To improve the compression ratio, gCom incorporates the following innovations. First, unlike other compression techniques that use frames or tiles as basic processing units, gCom is the first to employ a fine-grained processing unit (i.e., the color channel), enhancing repetition amplification without increasing raw data. Second, gCom introduces two key features— Hierarchical Delta and Channel Decorrelator —which maximize the locality of adjacent channels and reduce raw data size. Third, to maintain the original GPU throughput, gCom revolutionizes the Golomb-Rice algorithm and proposes a new compression approach, the Parallel-Oriented Golomb-Rice algorithm, enabling parallel execution of both decompression and compression processes. The entire design of gCom utilizes only idle resources and existing commands on mobile GPUs, thus keeping purchasing costs low. To date, gCom has improved the channel locality by nearly 50%. The best compression achievement received by gCom has reached around 20%. Dongjie Tang, Yun Wang 0039, Yicheng Gu, Fangxin Liu, Zhengwei Qi |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity VocoderabstractGenerative Adversarial Network (GAN) based vocoders are superior in inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator to promote GAN-based vocoders. Most existing time-frequency-representation-based discriminators are rooted in Short-Time Fourier Transform (STFT), whose time-frequency resolution in a spectrogram is fixed, making it incompatible with signals like singing voices that require flexible attention for different frequency bands. Motivated by that, our study utilizes the Constant-Q Transform (CQT), which owns dynamic resolution among frequencies, contributing to a better modeling ability in pitch accuracy and harmonic tracking. Specifically, we propose a Multi-Scale Sub-Band CQT (MS-SB-CQT) Discriminator, which operates on the CQT spectrogram at multiple scales and performs sub-band processing according to different octaves. Experiments conducted on both speech and singing voices confirm the effectiveness of our proposed method. Moreover, we also verified that the CQT-based and the STFT-based discriminators could be complementary under joint training. Specifically, enhanced by the proposed MS-SB-CQT and the existing MS-STFT Discriminators, the MOS of HiFi-GAN can be boosted from 3.27 to 3.87 for seen singers and from 3.40 to 3.78 for unseen singers. Yicheng Gu, Xueyao Zhang, Liumeng Xue, Zhizheng Wu 0001 |
ICASSP | 1 |
| 2024 | gHermes: Application-Unaware Acceleration for Cloud Rendering and Computing with Efficient GPU UtilizationabstractGPUs, as crucial tools for intensive computing tasks, are primarily used for rendering and computation.However, due to the high cost of GPUs and their frequent updates, owning powerful local GPUs is considered a luxury.Offloading tasks to the cloud is an effective solution.This method allows users to leverage the powerful capabilities of cloud-based high-performance computing resources to perform complex rendering tasks more quickly and cost-effectively.However, there has not yet been a solution that simultaneously considers both rendering and computation without requiring any modifications to local applications.This paper introduces gHermes, an application-unaware acceleration solution that automatically generates code to intercept rendering and computing tasks, overcoming the limitations of existing technologies.Additionally, for computational tasks, gHermes also enables fine-grained control of GPU utilization.Experimental results show that gHermes can handle both computing and rendering tasks efficiently.While ensuring the Quality of Service (QoS) requirements for rendering tasks (30-60 FPS), it effectively utilizes GPU resources for computation services. Xuyan Hu, Yicheng Gu, Zhengwei Qi |
SEKE | 2 |
| 2024 | Emilia: An Extensive, Multilingual, and Diverse Speech Dataset For Large-Scale Speech GenerationabstractRecent advancements in speech generation models have been significantly driven by the use of large-scale training data. However, producing highly spontaneous, human-like speech remains a challenge due to the scarcity of large, diverse, and spontaneous speech datasets. In response, we introduce Emilia, the first large-scale, multilingual, and diverse speech generation dataset. Emilia starts with over 101k hours of speech across six languages, covering a wide range of speaking styles to enable more natural and spontaneous speech generation. To facilitate the scale-up of Emilia, we also present Emilia-Pipe, the first open-source preprocessing pipeline designed to efficiently transform raw, in-the-wild speech data into high-quality training data with speech annotations. Experimental results demonstrate the effectiveness of both Emilia and Emilia-Pipe. Demos are available at: https://emilia-dataset.github.io/Emilia-Demo-Page/. Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu 0008, Jiaqi Li 0030, Peiyang Shi, Yuancheng Wang, Kai Chen 0026, Pengyuan Zhang, Zhizheng Wu 0001 |
SLT | 5 |
| 2024 | Leveraging Diverse Semantic-Based Audio Pretrained Models for Singing Voice ConversionabstractSinging Voice Conversion (SVC) is a technique that enables any singer to perform any song. To achieve this, it is essential to obtain speaker-agnostic representations from the source audio, which poses a significant challenge. A common solution involves utilizing a semantic-based audio pretrained model as a feature extractor However, the degree to which the extracted features can meet the SVC requirements remains an open question. This includes their capability to accurately model melody and lyrics, the speaker-independency of their underlying acoustic information, and their robustness for in-the-wild acoustic environments. In this study, we investigate the knowledge within classical semantic-based pretrained models in much detail. We discover that the knowledge of different models is diverse and can be complementary for SVC. Based on the above, we design a Singing Voice Conversion framework based on Diverse Semantic-based Feature Fusion (DSFF-SVC). Experimental results demonstrate that DSFF-SVC can be generalized and improve various existing SVC models, particularly in challenging real-world conversion tasks. Our demo website is available at https://diversesemanticsvc.github.io/. Xueyao Zhang, Zihao Fang, Yicheng Gu, Haopeng Chen, Lexiao Zou, Junan Zhang, Liumeng Xue, Zhizheng Wu 0001 |
SLT | 3 |
| 2024 | Amphion: an Open-Source Audio, Music, and Speech Generation ToolkitabstractAmphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion. Xueyao Zhang, Liumeng Xue, Yicheng Gu, Yuancheng Wang, Jiaqi Li 0030, Haorui He, Chaoren Wang, Songting Liu, Junan Zhang, Zihao Fang, Haopeng Chen, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Kai Chen 0026, Haizhou Li 0001, Zhizheng Wu 0001 |
SLT | 3 |
| 2024 | gVulkan: Scalable GPU Pooling for Pixel-Grained Rendering in Ray Tracing
Yicheng Gu, Yun Wang 0039, Yunfan Sun, Yuxin Xiang, Xuyan Hu, Zhengwei Qi, Haibing Guan |
USENIX ATC | 1 |
| 2024 | An Investigation of Time-Frequency Representation Discriminators for High-Fidelity VocodersabstractGenerative Adversarial Network (GAN) based vocoders are superior in both inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator for GAN-based vocoders. Most existing Time-Frequency Representation (TFR)-based discriminators are rooted in Short-Time Fourier Transform (STFT), which owns a constant Time-Frequency (TF) resolution, linearly scaled center frequencies, and a fixed decomposition basis, making it incompatible with signals like singing voices that require dynamic attention for different frequency bands and different time intervals. Motivated by that, we propose a Multi-Scale Sub-Band Constant-Q Transform CQT (MS-SB-CQT) discriminator and a Multi-Scale Temporal-Compressed Continuous Wavelet Transform CWT (MS-TC-CWT) discriminator. Both CQT and CWT have a dynamic TF resolution for different frequency bands. In contrast, CQT has a better modeling ability in pitch information, and CWT has a better modeling ability in short-time transients. Experiments conducted on both speech and singing voices confirm the effectiveness of our proposed discriminators. Moreover, the STFT, CQT, and CWT-based discriminators can be used jointly for better performance. The proposed discriminators can boost the synthesis quality of various state-of-the-art GAN-based vocoders, including HiFi-GAN, BigVGAN, and APNet. Yicheng Gu, Xueyao Zhang, Liumeng Xue, Haizhou Li 0001, Zhizheng Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | CARE: Cloudified Android With Optimized Rendering PlatformabstractDue to the excellent rendering capabilities, GPUs are mainstream accelerators in the Cloud-rendering industry. However, current Cloud-rendering systems suffer from a CPU-GPU workload imbalance that not only degrades application performance but also causes a significant waste of GPU resources. Recent proposals (such as API-forwarding and c-GPU) for improving CPU-GPU balance are promising but fail to solve system-resource redundancy issues (i.e., each instance tends to occupy all resources, exceeding its requirements). Such behavior will increase CPU load and lower effective GPU utilization. To demonstrate the severity of the issue, we evaluated real-world applications and results show that in most cases, nearly 50% of resources are useless. To solve this problem, we present CARE, the first framework intended to reduce the system-level redundancy by cloudifying the system from monolithic to Cloud-native. To allow users to configure required services, CARE puts forward a functional unit calledConfigurable Android (CA). To allow multiple instances to share certain types of resources, CARE innovatesSharing Resource (SR). To reduce the unused services, CARE introducesPruning Resources (PR). To further alleviate the CPU pressure and achieve CPU-GPU balance, we propose rShare, a system aiming at enhancing CPU effective utilization and increasing Android instance density of the Cloud-rendering platform. Based on Kubernetes, rShare divides all the CPUs into non-overlapping shared CPU pools, allocates instances to pools within milliseconds, and dynamically migrates them by tracking their QoS status. So far, CARE primarily focuses on Android systems and can handle 60 heavyweight instances (e.g., KOG (King of Glory)) on Intel SG1. rShare can apply instance allocation within milliseconds and increase the platform density by 39.4%. Yuxin Xiang, Dongjie Tang, Qiming Shi, Randy Xu, Mohammad R. Haghighat, Cathy Bao, Yicheng Gu, Zhengwei Qi, Haibing Guan |
IEEE Trans. Multim. | 10 |