Chenxi Gu

dblp:320/5862 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 90% Efficient and distributed learning · 10%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
watermarking
1.012026
SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking · ACL (1) 2026
Machine learning › Efficient and distributed learning › data-efficient learning
small-data learning
0.212022
Leveraging Similar Users for Personalized Language Modeling with Limited Data · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

sort-then-split by groups · 1.0KGW scheme · 1.0user similarity modeling · 0.6
YearPublicationVenuePosition
2026 SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking
abstract
Watermarking has emerged as a promising technique for tracing the authorship of content generated by large language models (LLMs).Among existing approaches, the KGW scheme is particularly attractive due to its versatility, efficiency, and effectiveness in natural language generation.However, KGW's effectiveness degrades significantly under low-entropy settings such as code generation and mathematical reasoning.A crucial step in the KGW method is random vocabulary partitioning, which enables adjustments to token selection based on specific preferences.Our study revealed that the nexttoken probability distribution plays an critical role in determining how much, or even whether, we can modify token selection and, consequently, the effectiveness of watermarking.We refer to this characteristic, associated with the probability distribution of each token prediction, as watermark strength.In cases of random vocabulary partitioning, the lower bound of watermark strength is dictated by the nexttoken probability distribution.However, we found that, by redesigning the vocabulary partitioning algorithm, we can potentially raise this lower bound.In this paper, we propose SSG (Sort-then-Split by Groups), a method that partitions the vocabulary into two logit-balanced subsets.This design lifts the lower bound of watermark strength for each token prediction, thereby improving watermark detectability.Experiments on code generation and mathematical reasoning datasets demonstrate the effectiveness of SSG.The source code is available at https://github.com/AllenG-L/SSG.
Chenxi Gu, Xiaoning Du 0001, John C. Grundy
ACL (1)1
2022 Leveraging Similar Users for Personalized Language Modeling with Limited Data
abstract
Charles Welch, Chenxi Gu, Jonathan Kummerfeld, Veronica Perez-Rosas, Rada Mihalcea. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea
ACL (1)2