Sunzhu Li

dblp:332/0754 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 38% Efficient and distributed learning · 30% Generative modeling · 13%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.122022
MorphTE: Injecting Morphology in Tensorized Embeddings · NeurIPS 2022
Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022
Natural language and speech › Language models and text generation
evaluation of language models
1.012026
RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation · ACL (1) 2026
Natural language and speech › Language models and text generation
instruction tuning
1.012026
RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation · ACL (1) 2026
Machine learning › Generative modeling › synthetic data generation
preference data synthesis
1.012026
RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model evaluation › automatic evaluation
rubric-based evaluation
1.012026
RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation · ACL (1) 2026
Natural language and speech › Machine translation
neural machine translation
0.612022
Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
tensor-train decomposition
0.612022
Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.612022
MorphTE: Injecting Morphology in Tensorized Embeddings · NeurIPS 2022
Machine learning › Efficient and distributed learning › model compression › embedding compression
word embedding compression
0.612022
MorphTE: Injecting Morphology in Tensorized Embeddings · NeurIPS 2022
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.212022
Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022
Machine learning › Deep learning architectures and training
transformer
0.212022
Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

coarse-to-fine generation · 1.0automated rubric generation · 1.0tensor-train decomposition · 0.6tensor product decomposition · 0.6morpheme vector · 0.6knowledge distillation · 0.6
YearPublicationVenuePosition
2026 RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation
abstract
Sunzhu Li, Jiale Zhao, Huimin Ren, Zhenlin Wei, Yang Zhou, Jingwen Yang, Shunyu Liu, Kaike Zhang, Chen Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sunzhu Li, Zhenlin Wei, Shunyu Liu 0001, Kaike Zhang
ACL (1)1
2022 Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation
abstract
Transformer has been demonstrated effective in Neural Machine Translation (NMT).However, it is memory-consuming and time-consuming in edge devices, resulting in some difficulties for real-time feedback.To compress and accelerate Transformer, we propose a Hybrid Tensor-Train (HTT) decomposition, which retains full rank and meanwhile reduces operations and parameters.A Transformer using HTT, named Hypoformer, consistently and notably outperforms the recent lightweight SOTA methods on three standard translation tasks under different parameter and speed scales.In extreme low resource scenarios, Hypoformer has a 7.1 point absolute improvement in BLEU and 1.27× speedup than the vanilla Transformer on the IWSLT'14 De-En task.
Sunzhu Li, Peng Zhang 0002, Guobing Gan, Xiuqing Lv, Benyou Wang, Victor Junqiu Wei, Xin Jiang 0002
EMNLP1
2022 MorphTE: Injecting Morphology in Tensorized Embeddings
abstract
In the era of deep learning, word embeddings are essential when dealing with text tasks. However, storing and accessing these embeddings requires a large amount of space. This is not conducive to the deployment of these models on resource-limited devices. Combining the powerful compression capability of tensor products, we propose a word embedding compression method with morphological augmentation, Morphologically-enhanced Tensorized Embeddings (MorphTE). A word consists of one or more morphemes, the smallest units that bear meaning or have a grammatical function. MorphTE represents a word embedding as an entangled form of its morpheme vectors via the tensor product, which injects prior semantic and grammatical knowledge into the learning of embeddings. Furthermore, the dimensionality of the morpheme vector and the number of morphemes are much smaller than those of words, which greatly reduces the parameters of the word embeddings. We conduct experiments on tasks such as machine translation and question answering. Experimental results on four translation datasets of different languages show that MorphTE can compress word embedding parameters by about $20$ times without performance loss and significantly outperforms related embedding compression methods.
Guobing Gan, Peng Zhang 0002, Sunzhu Li, Xiuqing Lu, Benyou Wang
NeurIPS3