VLDB 2026 Research / reviewers in the wild / expert
Sunzhu Li
dblp:332/0754
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 38% Efficient and distributed learning · 30% Generative modeling · 13% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
1.1 | 2 | 2022 | MorphTE: Injecting Morphology in Tensorized Embeddings · NeurIPS 2022 Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022 |
Natural language and speech › Language models and text generation
evaluation of language models |
1.0 | 1 | 2026 | RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation · ACL (1) 2026 |
Natural language and speech › Language models and text generation
instruction tuning |
1.0 | 1 | 2026 | RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation · ACL (1) 2026 |
Machine learning › Generative modeling › synthetic data generation
preference data synthesis |
1.0 | 1 | 2026 | RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation · ACL (1) 2026 |
Natural language and speech › Language models and text generation › large language model evaluation › automatic evaluation
rubric-based evaluation |
1.0 | 1 | 2026 | RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation · ACL (1) 2026 |
Natural language and speech › Machine translation
neural machine translation |
0.6 | 1 | 2022 | Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022 |
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
tensor-train decomposition |
0.6 | 1 | 2022 | Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022 |
Machine learning › Representation and self-supervised learning › word representation
word embedding |
0.6 | 1 | 2022 | MorphTE: Injecting Morphology in Tensorized Embeddings · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › model compression › embedding compression
word embedding compression |
0.6 | 1 | 2022 | MorphTE: Injecting Morphology in Tensorized Embeddings · NeurIPS 2022 |
Machine learning › Deep learning architectures and training › transformer
efficient transformer |
0.2 | 1 | 2022 | Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2022 | Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022 |
Methods — techniques the papers use, named apart from their topics
coarse-to-fine generation · 1.0automated rubric generation · 1.0tensor-train decomposition · 0.6tensor product decomposition · 0.6morpheme vector · 0.6knowledge distillation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine GenerationabstractSunzhu Li, Jiale Zhao, Huimin Ren, Zhenlin Wei, Yang Zhou, Jingwen Yang, Shunyu Liu, Kaike Zhang, Chen Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sunzhu Li, Zhenlin Wei, Shunyu Liu 0001, Kaike Zhang |
ACL (1) | 1 |
| 2022 | Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine TranslationabstractTransformer has been demonstrated effective in Neural Machine Translation (NMT).However, it is memory-consuming and time-consuming in edge devices, resulting in some difficulties for real-time feedback.To compress and accelerate Transformer, we propose a Hybrid Tensor-Train (HTT) decomposition, which retains full rank and meanwhile reduces operations and parameters.A Transformer using HTT, named Hypoformer, consistently and notably outperforms the recent lightweight SOTA methods on three standard translation tasks under different parameter and speed scales.In extreme low resource scenarios, Hypoformer has a 7.1 point absolute improvement in BLEU and 1.27× speedup than the vanilla Transformer on the IWSLT'14 De-En task. Sunzhu Li, Peng Zhang 0002, Guobing Gan, Xiuqing Lv, Benyou Wang, Victor Junqiu Wei, Xin Jiang 0002 |
EMNLP | 1 |
| 2022 | MorphTE: Injecting Morphology in Tensorized EmbeddingsabstractIn the era of deep learning, word embeddings are essential when dealing with text tasks. However, storing and accessing these embeddings requires a large amount of space. This is not conducive to the deployment of these models on resource-limited devices. Combining the powerful compression capability of tensor products, we propose a word embedding compression method with morphological augmentation, Morphologically-enhanced Tensorized Embeddings (MorphTE). A word consists of one or more morphemes, the smallest units that bear meaning or have a grammatical function. MorphTE represents a word embedding as an entangled form of its morpheme vectors via the tensor product, which injects prior semantic and grammatical knowledge into the learning of embeddings. Furthermore, the dimensionality of the morpheme vector and the number of morphemes are much smaller than those of words, which greatly reduces the parameters of the word embeddings. We conduct experiments on tasks such as machine translation and question answering. Experimental results on four translation datasets of different languages show that MorphTE can compress word embedding parameters by about $20$ times without performance loss and significantly outperforms related embedding compression methods. Guobing Gan, Peng Zhang 0002, Sunzhu Li, Xiuqing Lu, Benyou Wang |
NeurIPS | 3 |