Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Hongkun Hao

dblp:349/2933 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0006-4270-8455ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 85% Generative modeling · 15%
Network and information security
1 paper
Digital forensics and information hiding · 100%
Computer graphics and multimedia
1 paper
Image and video coding · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
decoding
1.422024
Improving Open-Ended Text Generation via Adaptive Decoding · ICML 2024
Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text Generation · EMNLP 2023
Natural language and speech › Language models and text generation › decoding
adaptive decoding
0.812024
Improving Open-Ended Text Generation via Adaptive Decoding · ICML 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
G-Refine: A General Quality Refiner for Text-to-Image Generation · ACM Multimedia 2024
Digital forensics and information hiding › watermarking
text watermarking
0.812024
Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation › text generation
open-ended text generation
0.712023
Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text Generation · EMNLP 2023
Image and video coding
image quality assessment
0.212024
G-Refine: A General Quality Refiner for Text-to-Image Generation · ACM Multimedia 2024

Methods — techniques the papers use, named apart from their topics

quality refinement · 1.5machine translation · 1.5entropy-based confidence · 0.8candidate set selection · 0.8penalty decoding · 0.7length penalty · 0.7forgetting mechanism · 0.7
YearPublicationVenuePosition
2025 Boosting Large Language Model for Speech Synthesis: An Empirical Study
abstract
Large language models (LLMs) have made significant advancements in natural language processing and are concurrently extending the language ability to other modalities, such as speech and vision. Nevertheless, most of the previous work focuses on prompting LLMs with perception abilities like auditory comprehension, and the effective approach for augmenting LLMs with speech synthesis capabilities remains ambiguous. In this paper, we conduct a comprehensive empirical exploration of boosting LLMs with the ability to generate speech, by combining pre-trained LLM LLaMA/OPT and text-to-speech synthesis model VALL-E. We compare three integration methods between LLMs and speech synthesis models, including directly fine-tuned LLMs, superposed layers of LLMs and VALL-E, and coupled LLMs and VALL-E using LLMs as a powerful text encoder. Experimental results show that, using LoRA method to fine-tune LLMs directly to boost the speech synthesis capability does not work well, and superposed LLMs and VALL-E can improve the quality of generated speech both in speaker similarity and word error rate (WER). Among these three methods, coupled methods leveraging LLMs as the text encoder can achieve the best performance, making it outperform original speech synthesis models with a consistently better speaker similarity and a significant (10.9%) WER reduction.
Hongkun Hao, Shujie Liu 0001, Jinyu Li 0001, Shujie Hu, Rui Wang 0015, Furu Wei
ICASSP1
2024 Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models
abstract
Zhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang, Zhaopeng Tu, Zhuosheng Zhang, Rui Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhiwei He 0002, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang 0007, Zhaopeng Tu, Zhuosheng Zhang 0001, Rui Wang 0015
ACL (1)3
2024 Q-Refine: A Perceptual Quality Refiner for AI-Generated Image
abstract
With the rapid evolution of the Text-to-Image (T2I) model in recent years, their unsatisfactory generation result has become a challenge. However, uniformly refining AI-Generated Images (AIGIs) of different qualities not only limited optimization capabilities for low-quality AIGIs but also brought negative optimization to high-quality AIGIs. To address this issue, a quality-award refiner named Q-Refine is proposed. Based on the preference of the Human Visual System (HVS), Q-Refine uses the Image Quality Assessment (IQA) metric to guide the refining process for the first time, and modify images of different qualities through three adaptive pipelines. Experimental data shows that for mainstream T2I models, Q-Refine can perform effective optimization to AIGIs of different qualities. It can be a general refiner to optimize AIGIs from both fidelity and aesthetic quality levels, thus expanding the application of the T2I generation models. The code is released on https://github.com/Q-Future/Q-Refine.
Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Kaiwei Zhang, Lei Bai 0001, Xiaohong Liu 0001, Xiongkuo Min, Weisi Lin, Guangtao Zhai
ICME4
2024 Improving Open-Ended Text Generation via Adaptive Decoding
abstract
Current language models decode text token by token according to probabilistic distribution, and determining the appropriate candidates for the next token is crucial to ensure generation quality. This study introduces adaptive decoding, a mechanism that dynamically empowers language models to ascertain a sensible candidate set during generation. Specifically, we introduce an entropy-based metric called confidence and conceptualize determining the optimal candidate set as a confidence-increasing process. The rationality of including a token in the candidate set is assessed by leveraging the increment of confidence. Experimental results reveal that our method balances diversity and coherence well. The human evaluation shows that our method can generate human-preferred text. Additionally, our method can potentially improve the reasoning ability of language models.
Hongkun Hao, Zhiwei He 0002, Yiming Ai, Rui Wang 0015
ICML2
2024 G-Refine: A General Quality Refiner for Text-to-Image Generation
Chunyi Li 0001, Haoning Wu 0001, Hongkun Hao, Tengchuan Kou, Chaofeng Chen, Lei Bai 0001, Xiaohong Liu 0001, Weisi Lin, Guangtao Zhai
ACM Multimedia3
2023 Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text Generation
abstract
The decoding algorithm is critical for openended text generation, transforming latent representations into coherent and meaningful outputs.This paper investigates the selfreinforcement effect in text generation and the effectiveness of a repetition penalty to mitigate it.However, determining the optimal repetition penalty value is challenging.To tackle this, we propose a forgetting mechanism that disregards distant tokens, reducing the burden of penalty selection.In addition, we introduce a length penalty to address overly short sentences caused by excessive penalties.Our penalty decoding approach incorporating three strategies helps resolve issues with sampling methods deviating from factual information.Experimental results demonstrate the efficacy of our approach in generating high-quality sentences resembling human output. 1
Hongkun Hao, Rui Wang 0015
EMNLP2