EDBT 2026 Demo / reviewers in the wild / expert
Shengding Hu
dblp:268/5534
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2025
0009-0000-8037-2055ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 15 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language ModelsabstractActivation sparsity refers to the existence of considerable weakly-contributed elements among activation outputs, serving as a promising paradigm for accelerating model inference. Nevertheless, most large language models (LLMs) adopt activation functions without intrinsic activation sparsity (e.g., GELU and Swish). Some recent efforts have explored introducing ReLU or its variants as the substitutive activation function to pursue activation sparsity and acceleration, but few can simultaneously obtain high activation sparsity and comparable model performance. This paper introduces a simple and effective method named “ProSparse” to sparsify LLMs while achieving both targets. Specifically, after introducing ReLU activation, ProSparse adopts progressive sparsity regularization with a factor smoothly increasing for multiple stages. This can enhance activation sparsity and mitigate performance degradation by avoiding radical shifts in activation distributions. With ProSparse, we obtain high sparsity of 89.32% for LLaMA2-7B, 88.80% for LLaMA2-13B, and 87.89% for end-size MiniCPM-1B, respectively, with comparable performance to their original Swish-activated versions. These present the most sparsely activated models among open-source LLaMA versions and competitive end-size models. Inference acceleration experiments further demonstrate the significant practical acceleration potential of LLMs with higher activation sparsity, obtaining up to 4.52x inference speedup. Xu Han 0007, Zhengyan Zhang, Shengding Hu, Xiyu Shi, Kuai Li, Zhiyuan Liu 0001, Guangli Li, Maosong Sun 0001 |
COLING | 4 |
| 2025 | A Multi-Power Law for Loss Curve Prediction Across Learning Rate SchedulesabstractTraining large models is both resource-intensive and time-consuming, making it crucial to understand the quantitative relationship between model performance and hyperparameters. In this paper, we derive an empirical law that predicts pretraining loss for large language models for every intermediate training step across various learning rate schedules, including constant, cosine, and step decay schedules. Our proposed law takes a multi-power form, combining a power law based on the sum of learning rates and additional power laws to account for a loss reduction effect as learning rate decays. We validate this law extensively on Llama-2 models of varying sizes and demonstrate that, after fitting on a few learning rate schedules, it accurately predicts the loss curves for unseen schedules of different shapes and horizons. Moreover, by minimizing the predicted final pretraining loss across learning rate schedules, we are able to find a schedule that outperforms the widely-used cosine learning rate schedule. Interestingly, this automatically discovered schedule bears some resemblance to the recently proposed Warmup-Stable-Decay (WSD) schedule (Hu et al, 2024) but achieves a slightly lower final loss. We believe these results could offer valuable insights for understanding the dynamics of pretraining and for designing learning rate schedules to improve efficiency. Kairong Luo, Haodong Wen, Shengding Hu, Zhenbo Sun, Zhiyuan Liu 0001, Maosong Sun 0001, Kaifeng Lyu |
ICLR | 3 |
| 2024 | OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsabstractChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Thai, Junhao Shen, Jinyi Hu, Xu Han, Yujie Huang, Yuxiang Zhang, Jie Liu, Lei Qi, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu, Zhen Leng Thai, Jinyi Hu, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 4 |
| 2024 | ınftyBench: Extending Long Context Evaluation Beyond 100K TokensabstractXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, Junhao Chen, Moo Hao, Xu Han, Zhen Thai, Shuo Wang, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yingfa Chen, Shengding Hu, Zihang Xu, Moo Khai Hao, Xu Han 0007, Zhen Leng Thai, Shuo Wang 0013, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 3 |
| 2024 | Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex ModelsabstractXinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun, Zhiyuan Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yingfa Chen, Shengding Hu, Xu Han 0007, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun 0001, Zhiyuan Liu 0001 |
EMNLP | 3 |
| 2024 | DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language ModelsabstractRanchi Zhao, Zhen Leng Thai, Yifan Zhang, Shengding Hu, Jie Zhou, Yunqi Ba, Jie Cai, Zhiyuan Liu, Maosong Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ranchi Zhao, Zhen Leng Thai, Shengding Hu, Jie Zhou 0024, Yunqi Ba, Jie Cai 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 4 |
| 2024 | Predicting Emergent Abilities with Infinite Resolution EvaluationabstractThe scientific scale-up of large language models (LLMs) necessitates a comprehensive understanding of their scaling properties. However, the existing literature on the scaling properties only yields an incomplete answer: optimization loss decreases predictably as the model size increases, in line with established scaling law; yet no scaling law for task has been established and the task performances are far from predictable during scaling. Task performances typically show minor gains on small models until they improve dramatically once models exceed a size threshold, exemplifying the ''emergent abilities''. In this study, we discover that small models, although they exhibit minor performance, demonstrate critical and consistent task performance improvements that are not captured by conventional evaluation strategies due to insufficient measurement resolution. To measure such improvements, we introduce PassUntil, an evaluation strategy with theoretically infinite resolution, through massive sampling in the decoding phase. With PassUntil, we conduct a quantitative investigation into the scaling law of task performance. The investigation contains two parts. Firstly, a strict task scaling law that is not conventionally known to exist, is identified, enhancing the predictability of task performances. Remarkably, we are able to predict the performance of the 2.4B model on code generation with merely 0.05\% deviation before training starts, which is the first systematic attempt to verify predictable scaling proposed by GPT-4's report. Secondly, underpinned by PassUntil, we are able to study emergent abilities quantitatively. We identify a kind of accelerated emergence whose scaling curve cannot be fitted by standard scaling law function and has a increasing speed. We then examine two hypothesis and imply that the ``multiple circuits hypothesis'' might be responsible for the accelerated emergence. Shengding Hu, Xin Liu 0086, Xu Han 0007, Chaoqun He, Weilin Zhao, Yankai Lin 0001, Ning Ding 0002, Zebin Ou, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 1 |
| 2023 | Exploring Lottery Prompts for Pre-trained Language ModelsabstractYulin Chen, Ning Ding, Xiaobin Wang, Shengding Hu, Haitao Zheng, Zhiyuan Liu, Pengjun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yulin Chen 0001, Ning Ding 0002, Xiaobin Wang, Shengding Hu, Hai-Tao Zheng 0002, Zhiyuan Liu 0001, Pengjun Xie |
ACL (1) | 4 |
| 2023 | Won't Get Fooled Again: Answering Questions with False PremisesabstractPre-trained language models (PLMs) have shown unprecedented potential in various fields, especially as the backbones for questionanswering (QA) systems.However, they tend to be easily deceived by tricky questions such as "How many eyes does the sun have?".Such frailties of PLMs often allude to the lack of knowledge within them.In this paper, we find that the PLMs already possess the knowledge required to rebut such questions, and the key is how to activate the knowledge.To systematize this observation, we investigate the PLMs' responses to one kind of tricky questions, i.e., the false premises questions (FPQs).We annotate a FalseQA dataset containing 2365 human-written FPQs, with the corresponding explanations for the false premises and the revised true premise questions.Using FalseQA, we discover that PLMs are capable of discriminating FPQs by fine-tuning on moderate numbers (e.g., 256) of examples.PLMs also generate reasonable explanations for the false premise, which serve as rebuttals.Further replaying a few general questions during training allows PLMs to excel on FPQs and general questions simultaneously.Our work suggests that once the rebuttal ability is stimulated, knowledge inside the PLMs can be effectively utilized to handle FPQs, which incentivizes the research on PLM-based QA systems.The FalseQA dataset and code are available at https://github.com/thunlp/FalseQA. Shengding Hu, Xingyi Cheng, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 1 |
| 2023 | Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsabstractFine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT.Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading to improved performance.This paper aims to push the upper bound of opensource models further.We first provide a systematically designed, diverse, informative, large-scale dataset of instructional conversations, UltraChat, which does not involve human queries.Our objective is to capture the breadth of interactions between a human user and an AI assistant and employs a comprehensive framework to generate multi-turn conversation iteratively.UltraChat contains 1.5 million high-quality multi-turn dialogues and covers a wide range of topics and instructions.Our statistical analysis of UltraChat reveals its superiority in various key metrics, including scale, average length, diversity, coherence, etc., solidifying its position as a leading opensource dataset.Building upon UltraChat, we fine-tune a LLaMA model to create a powerful conversational model, UltraLM.Our evaluations indicate that UltraLM consistently outperforms other open-source models, including WizardLM and Vicuna, the previously recognized state-of-the-art open-source models. Ning Ding 0002, Yulin Chen 0001, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu 0001, Maosong Sun 0001, Bowen Zhou 0002 |
EMNLP | 5 |
| 2023 | Exploring the Impact of Model Scaling on Parameter-Efficient TuningabstractYusheng Su, Chi-Min Chan, Jiali Cheng, Yujia Qin, Yankai Lin, Shengding Hu, Zonghan Yang, Ning Ding, Xingzhi Sun, Guotong Xie, Zhiyuan Liu, Maosong Sun. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yusheng Su, Chi-Min Chan, Jiali Cheng, Yujia Qin, Yankai Lin 0001, Shengding Hu, Zonghan Yang, Ning Ding 0002, Xingzhi Sun 0002, Guo Tong Xie, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 6 |
| 2022 | Prototypical Verbalizer for Prompt-based Few-shot TuningabstractPrompt-based tuning for pre-trained language models (PLMs) has shown its effectiveness in few-shot learning.Typically, prompt-based tuning wraps the input text into a cloze question.To make predictions, the model maps the output words to labels via a verbalizer, which is either manually designed or automatically built.However, manual verbalizers heavily depend on domain-specific prior knowledge and human efforts, while finding appropriate label words automatically still remains challenging.In this work, we propose the prototypical verbalizer (ProtoVerb) which is built directly from training data.Specifically, Pro-toVerb learns prototype vectors as verbalizers by contrastive learning.In this way, the prototypes summarize training instances and are able to enclose rich class-level semantics.We conduct experiments on both topic classification and entity typing tasks, and the results demonstrate that ProtoVerb significantly outperforms current automatic verbalizers, especially when training data is extremely scarce.More surprisingly, ProtoVerb consistently boosts promptbased tuning even on untuned PLMs, indicating an elegant non-tuning way to utilize PLMs.Our codes are avaliable at https: //github.com/thunlp/OpenPrompt. Ganqu Cui, Shengding Hu, Ning Ding 0002, Longtao Huang, Zhiyuan Liu 0001 |
ACL (1) | 2 |
| 2022 | Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text ClassificationabstractShengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu, Jingang Wang, Juanzi Li, Wei Wu, Maosong Sun. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shengding Hu, Ning Ding 0002, Zhiyuan Liu 0001, Jingang Wang, Juan-Zi Li, Wei Wu 0014, Maosong Sun 0001 |
ACL (1) | 1 |
| 2022 | COPEN: Probing Conceptual Knowledge in Pre-trained Language ModelsabstractConceptual knowledge is fundamental to human cognition and knowledge bases.However, existing knowledge probing works only focus on evaluating factual knowledge of pre-trained language models (PLMs) and ignore conceptual knowledge.Since conceptual knowledge often appears as implicit commonsense behind texts, designing probes for conceptual knowledge is hard.Inspired by knowledge representation schemata, we comprehensively evaluate conceptual knowledge of PLMs by designing three tasks to probe whether PLMs organize entities by conceptual similarities, learn conceptual properties, and conceptualize entities in contexts, respectively.For the tasks, we collect and annotate 24k data instances covering 393 concepts, which is COPEN, a COnceptual knowledge Probing bENchmark.Extensive experiments on different sizes and types of PLMs show that existing PLMs systematically lack conceptual knowledge and suffer from various spurious correlations.We believe this is a critical bottleneck for realizing human-like cognition in PLMs.COPEN and our codes are publicly released at https: //github.com/THU-KEG/COPEN. Hao Peng 0015, Xiaozhi Wang, Shengding Hu, Hailong Jin, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Qun Liu 0001 |
EMNLP | 3 |
| 2022 | Sparse Structure Search for Delta TuningabstractAdapting large pre-trained models (PTMs) through fine-tuning imposes prohibitive computational and storage burdens. Recent studies of delta tuning (DT), i.e., parameter-efficient tuning, find that only optimizing a small portion of parameters conditioned on PTMs could yield on-par performance compared to conventional fine-tuning. Generally, DT methods exquisitely design delta modules (DT modules) which could be applied to arbitrary fine-grained positions inside PTMs. However, the effectiveness of these fine-grained positions largely relies on sophisticated manual designation, thereby usually producing sub-optimal results. In contrast to the manual designation, we explore constructing DT modules in an automatic manner. We automatically \textbf{S}earch for the \textbf{S}parse \textbf{S}tructure of \textbf{Delta} Tuning (S$^3$Delta). Based on a unified framework of various DT methods, S$^3$Delta conducts the differentiable DT structure search through bi-level optimization and proposes shifted global sigmoid method to explicitly control the number of trainable parameters. Extensive experiments show that S$^3$Delta surpasses manual and random structures with less trainable parameters. The searched structures preserve more than 99\% fine-tuning performance with 0.01\% trainable parameters. Moreover, the advantage of S$^3$Delta is amplified with extremely low trainable parameters budgets (0.0009\%$\sim$0.01\%). The searched structures are transferable and explainable, providing suggestions and guidance for the future design of DT methods. Our codes are publicly available at \url{https://github.com/thunlp/S3Delta}. Shengding Hu, Zhen Zhang 0008, Ning Ding 0002, Yadao Wang, Yasheng Wang, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 1 |
| 2020 | Graph Policy Network for Transferable Active Learning on GraphsabstractGraph neural networks (GNNs) have been attracting increasing popularity due to their simplicity and effectiveness in a variety of fields. However, a large number of labeled data is generally required to train these networks, which could be very expensive to obtain in some domains. In this paper, we study active learning for GNNs, i.e., how to efficiently label the nodes on a graph to reduce the annotation cost of training GNNs. We formulate the problem as a sequential decision process on graphs and train a GNN-based policy network with reinforcement learning to learn the optimal query strategy. By jointly training on several source graphs with full labels, we learn a transferable active learning policy which can directly generalize to unlabeled target graphs. Experimental results on multiple datasets from different domains prove the effectiveness of the learned policy in promoting active learning performance in both settings of transferring between graphs in the same domain and across different domains. Shengding Hu, Zheng Xiong, Meng Qu, Xingdi Yuan, Marc-Alexandre Côté, Zhiyuan Liu 0001, Jian Tang 0005 |
NeurIPS | 1 |