Zhujin Gao

dblp:336/4920 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 83% Machine translation · 17%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.522025
Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion Processes · ACL (1) 2025
DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation · EMNLP 2023
Machine learning › Generative modeling › diffusion model › score-based generative model
continuous-time diffusion
0.912025
Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion Processes · ACL (1) 2025
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model
0.912025
Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion Processes · ACL (1) 2025
Natural language and speech › Machine translation › speech translation
speech-to-speech translation
0.712023
DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

time predictor · 0.9poisson diffusion process · 0.9noise schedule optimization · 0.9discrete speech units · 0.7diffusion model · 0.7
YearPublicationVenuePosition
2025 Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion Processes
abstract
Diffusion models have emerged as a promising approach for text generation, with recent works falling into two main categories: discrete and continuous diffusion models.Discrete diffusion models apply token corruption independently using categorical distributions, allowing for different diffusion progress across tokens but lacking fine-grained control.Continuous diffusion models map tokens to continuous spaces and apply fine-grained noise, but the diffusion progress is uniform across tokens, limiting their ability to capture semantic nuances.To address these limitations, we propose Non-simultaneous Continuous Diffusion Models (NeoDiff), a novel diffusion model that integrates the strengths of both discrete and continuous approaches.NeoDiff introduces a Poisson diffusion process for the forward process, enabling a flexible and fine-grained noising paradigm, and employs a time predictor for the reverse process to adaptively modulate the denoising progress based on token semantics.Furthermore, NeoDiff utilizes an optimized schedule for inference to ensure more precise noise control and improved performance.Our approach unifies the theories of discrete and continuous diffusion models, offering a more principled and effective framework for text generation.Experimental results on several text generation tasks demonstrate NeoDiff's superior performance compared to baselines of nonautoregressive continuous and discrete diffusion models, iterative-based methods and autoregressive diffusion-based methods.These results highlight NeoDiff's potential as a powerful tool for generating high-quality text and advancing the field of diffusion-based text generation.
Bocheng Li, Zhujin Gao
ACL (1)2
2024 Few-shot Temporal Pruning Accelerates Diffusion Models for Text Generation
abstract
Diffusion models have achieved significant success in computer vision and shown immense potential in natural language processing applications, particularly for text generation tasks. However, generating high-quality text using these models often necessitates thousands of iterations, leading to slow sampling rates. Existing acceleration methods either neglect the importance of the distribution of sampling steps, resulting in compromised performance with smaller number of iterations, or require additional training, introducing considerable computational overheads. In this paper, we present Few-shot Temporal Pruning, a novel technique designed to accelerate diffusion models for text generation without supplementary training while effectively leveraging limited data. Employing a Bayesian optimization approach, our method effectively eliminates redundant sampling steps during the sampling process, thereby enhancing the generation speed. A comprehensive evaluation of discrete and continuous diffusion models across various tasks, including machine translation, question generation, and paraphrasing, reveals that our approach achieves competitive performance even with minimal sampling steps after down to less than 1 minute of optimization, yielding a significant acceleration of up to 400x in text generation tasks.
Bocheng Li, Zhujin Gao, Yongxin Zhu 0003, Kun Yin, Haoyu Cao 0001, Deqiang Jiang, Linli Xu 0002
LREC/COLING2
2024 Empowering Diffusion Models on the Embedding Space for Text Generation
abstract
Zhujin Gao, Junliang Guo, Xu Tan, Yongxin Zhu, Fang Zhang, Jiang Bian, Linli Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhujin Gao, Junliang Guo, Xu Tan 0003, Yongxin Zhu 0003, Fang Zhang 0006, Jiang Bian 0002, Linli Xu 0002
NAACL-HLT1
2023 DiffS2UT: A Semantic Preserving Diffusion Model for Textless Direct Speech-to-Speech Translation
abstract
While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial problem.Specifically, due to the low information density of speech data, the transformed discrete speech unit sequence is much longer than the corresponding text transcription, posing significant challenges to existing auto-regressive models.Furthermore, it is not optimal to brutally apply discrete diffusion on the speech unit sequence while disregarding the continuous space structure, which will degrade the generation performance significantly.In this paper, we propose a novel diffusion model by applying the diffusion forward process in the continuous speech representation space, while employing the diffusion backward process in the discrete speech unit space.In this way, we preserve the semantic structure of the continuous speech representation space in the diffusion process and integrate the continuous and discrete diffusion models.We conduct extensive experiments on the textless direct speech-to-speech translation task, where the proposed method achieves comparable results to the computationally intensive auto-regressive baselines (500 steps on average) with significantly fewer decoding steps (50 steps).
Yongxin Zhu 0003, Zhujin Gao, Xinyuan Zhou, Zhongyi Ye, Linli Xu 0002
EMNLP2
2023 ItrievalKD: An Iterative Retrieval Framework Assisted with Knowledge Distillation for Noisy Text-to-Image Retrieval
Yongxin Zhu 0003, Zhujin Gao, Xin Sheng 0003, Linli Xu 0002
PAKDD (3)3