VLDB 2026 Research / reviewers in the wild / expert
Zhiyuan Zhao 0001
dblp:93/5901-1
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MicroCinema: A Divide-and-Conquer Approach for Text-to-Video GenerationabstractWe present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer strategy which divides the text-to-video into a two-stage process: text-to-image generation and image&text-to-video generation. This strategy offers two significant advantages. a) It allows us to take full advantage of the recent advances in text-to-image models, such as Stable Diffusion, Midjourney, and DALLE, to generate photorealistic and highly detailed images. b) Leveraging the generated image, the model can allocate less focus to fine-grained appearance details, prioritizing the efficient learning of motion dynamics. To implement this strategy effectively, we introduce two core designs. First, we propose the Appearance Injection Network, enhancing the preservation of the appearance of the given image. Second, we introduce the appearance Noise Prior, a novel mechanism aimed at maintaining the capabilities of pre-trained 2D diffusion models. These design elements empower MicroCinema to generate high-quality videos with precise motion, guided by the provided text prompts. Extensive experiments demonstrate the superiority of the proposed framework. Concretely, MicroCinema achieves SOTA zero-shot FVD of 342.86 on UCF-JOJ and 377.40 on MSR-VTT. Jianmin Bao, Wenming Weng, Ruoyu Feng 0001, Dacheng Yin, Jingxu Zhang, Qi Dai 0001, Zhiyuan Zhao 0001, Chunyu Wang 0001, Yuhui Yuan, Xiaoyan Sun 0001, Chong Luo 0001, Baining Guo |
CVPR | 9 |
| 2023 | Filler Word Detection with Hard Category Mining and Inter-Category Focal LossabstractFiller words like "um" or "uh" are common in spontaneous speech. It is desirable to automatically detect and remove them in recordings, as they affect the fluency, confidence, and professionalism of speech. Previous studies and our preliminary experiments reveal that the biggest challenge in filler word detection is that fillers can be easily confused with other hard categories like "a" or "I". In this paper, we propose a novel filler word detection method that effectively addresses this challenge by adding auxiliary categories dynamically and applying an additional inter-category focal loss. The auxiliary categories force the model to explicitly model the confusing words by mining hard categories. In addition, inter-category focal loss adaptively adjusts the penalty weight between "filler" and "non-filler" categories to deal with other confusing words left in the "non-filler" category. Our system achieves the best results, with a huge improvement compared to other methods on the PodcastFillers dataset. Zhiyuan Zhao 0001, Chuanxin Tang, Dacheng Yin, Chong Luo 0001 |
ICASSP | 1 |
| 2023 | TridentSE: Guiding Speech Enhancement with 32 Global Tokens
Dacheng Yin, Zhiyuan Zhao 0001, Chuanxin Tang, Zhiwei Xiong, Chong Luo 0001 |
INTERSPEECH | 2 |
| 2022 | RetrieverTTS: Modeling Decomposed Factors for Text-Based Speech InsertionabstractThis paper proposes a new "decompose-and-edit" paradigm for the text-based speech insertion task that facilitates arbitrarylength speech insertion and even full sentence generation.In the proposed paradigm, global and local factors in speech are explicitly decomposed and separately manipulated to achieve high speaker similarity and continuous prosody.Specifically, we proposed to represent the global factors by multiple tokens, which are extracted by cross-attention operation and then injected back by link-attention operation.Due to the rich representation of global factors, we manage to achieve high speaker similarity in a zero-shot manner.In addition, we introduce a prosody smoothing task to make the local prosody factor context-aware and therefore achieve satisfactory prosody continuity.We further achieve high voice quality with an adversarial training stage.In the subjective test, our method achieves state-of-the-art performance in both naturalness and similarity.Audio samples can be found at https://ydcustc.github.io/retrieverTTS-demo/. Dacheng Yin, Chuanxin Tang, Xiaoqiang Wang 0006, Zhiyuan Zhao 0001, Zhiwei Xiong, Sheng Zhao 0002, Chong Luo 0001 |
INTERSPEECH | 5 |
| 2022 | An Anchor-Free Detector for Continuous Speech Keyword SpottingabstractContinuous Speech Keyword Spotting (CSKWS) is a task to detect predefined keywords in a continuous speech.In this paper, we regard CSKWS as a one-dimensional object detection task and propose a novel anchor-free detector, named AF-KWS, to solve the problem.AF-KWS directly regresses the center locations and lengths of the keywords through a single-stage deep neural network.In particular, AF-KWS is tailored for this speech task as we introduce an auxiliary unknown class to exclude other words from non-speech or silent background.We have built two benchmark datasets named LibriTop-20 and continuous meeting analysis keywords (CMAK) dataset for CSKWS.Evaluations on these two datasets show that our proposed AF-KWS outperforms reference schemes by a large margin, and therefore provides a decent baseline for future research. Zhiyuan Zhao 0001, Chuanxin Tang, Chengdong Yao, Chong Luo 0001 |
INTERSPEECH | 1 |
| 2021 | Zero-Shot Text-to-Speech for Text-Based Insertion in Audio NarrationabstractGiven a piece of speech and its transcript text, text-based speech editing aims to generate speech that can be seamlessly inserted into the given speech by editing the transcript.Existing methods adopt a two-stage approach: synthesize the input text using a generic text-to-speech (TTS) engine and then transform the voice to the desired voice using voice conversion (VC).A major problem of this framework is that VC is a challenging problem which usually needs a moderate amount of parallel training data to work satisfactorily.In this paper, we propose a one-stage context-aware framework to generate natural and coherent target speech without any training data of the target speaker.In particular, we manage to perform accurate zero-shot duration prediction for the inserted text.The predicted duration is used to regulate both text embedding and speech embedding.Then, based on the aligned cross-modality input, we directly generate the mel-spectrogram of the edited speech with a transformer-based decoder.Subjective listening tests show that despite the lack of training data for the speaker, our method has achieved satisfactory results.It outperforms a recent zero-shot TTS engine by a large margin. Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Dacheng Yin, Wenjun Zeng 0001 |
Interspeech | 3 |
| 2020 | Joint Time-Frequency and Time Domain Learning for Speech EnhancementabstractFor single-channel speech enhancement, both time-domain and time-frequency-domain methods have their respective pros and cons. In this paper, we present a cross-domain framework named TFT-Net, which takes time-frequency spectrogram as input and produces time-domain waveform as output. Such a framework takes advantage of the knowledge we have about spectrogram and avoids some of the drawbacks that T-F-domain methods have been suffering from. In TFT-Net, we design an innovative dual-path attention block (DAB) to fully exploit correlations along the time and frequency axes. We further discover that a sample-independent DAB (SDAB) achieves a good tradeoff between enhanced speech quality and complexity. Ablation studies show that both the cross-domain design and the SDAB block bring large performance gain. When logarithmic MSE is used as the training criteria, TFT-Net achieves the highest SDR and SSNR among state-of-the-art methods on two major speech enhancement benchmarks. Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Wenxuan Xie, Wenjun Zeng 0001 |
IJCAI | 3 |