VLDB 2026 Research / reviewers in the wild / expert
Chenxi Gu
dblp:320/5862
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 90% Efficient and distributed learning · 10% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
watermarking |
1.0 | 1 | 2026 | SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking · ACL (1) 2026 |
Machine learning › Efficient and distributed learning › data-efficient learning
small-data learning |
0.2 | 1 | 2022 | Leveraging Similar Users for Personalized Language Modeling with Limited Data · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
sort-then-split by groups · 1.0KGW scheme · 1.0user similarity modeling · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSG: Logit-Balanced Vocabulary Partitioning for LLM WatermarkingabstractWatermarking has emerged as a promising technique for tracing the authorship of content generated by large language models (LLMs).Among existing approaches, the KGW scheme is particularly attractive due to its versatility, efficiency, and effectiveness in natural language generation.However, KGW's effectiveness degrades significantly under low-entropy settings such as code generation and mathematical reasoning.A crucial step in the KGW method is random vocabulary partitioning, which enables adjustments to token selection based on specific preferences.Our study revealed that the nexttoken probability distribution plays an critical role in determining how much, or even whether, we can modify token selection and, consequently, the effectiveness of watermarking.We refer to this characteristic, associated with the probability distribution of each token prediction, as watermark strength.In cases of random vocabulary partitioning, the lower bound of watermark strength is dictated by the nexttoken probability distribution.However, we found that, by redesigning the vocabulary partitioning algorithm, we can potentially raise this lower bound.In this paper, we propose SSG (Sort-then-Split by Groups), a method that partitions the vocabulary into two logit-balanced subsets.This design lifts the lower bound of watermark strength for each token prediction, thereby improving watermark detectability.Experiments on code generation and mathematical reasoning datasets demonstrate the effectiveness of SSG.The source code is available at https://github.com/AllenG-L/SSG. Chenxi Gu, Xiaoning Du 0001, John C. Grundy |
ACL (1) | 1 |
| 2022 | Leveraging Similar Users for Personalized Language Modeling with Limited DataabstractCharles Welch, Chenxi Gu, Jonathan Kummerfeld, Veronica Perez-Rosas, Rada Mihalcea. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, Rada Mihalcea |
ACL (1) | 2 |