Damai Dai

dblp:199/2097 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
22since 2021 · last 2026
0009-0004-9714-7902ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Large Language Models Struggle with Unreasonability in Math Problems
abstract
Large Language Models (LLMs) have shown remarkable success on a wide range of math and reasoning benchmarks. However, we observe that they often struggle when faced with unreasonable math problems. Instead of recognizing these issues, models frequently proceed as if the problem is well-posed, producing incorrect answers or falling into overthinking and verbose self-correction. To systematically investigate this overlooked vulnerability, we propose the Unreasonable Math Problems (UMP) benchmark, designed to evaluate LLMs' ability to detect and respond to unreasonable math problem statements. Based on extensive experiments covering 19 LLMs, we find that even state-of-the-art general models like GPT-4o struggle on UMP. While reasoning models such as DeepSeek-R1 demonstrate a higher sensitivity to unreasonable inputs, this often comes at the cost of generating overly long and meaningless responses that fail to converge. We further find that prompting and fine-tuning enhance the detection of unreasonable inputs, with minor and acceptable trade-offs, making them practical solutions in this challenging setting.
Jingyuan Ma, Damai Dai, Zihang Yuan, Rui Li 0094, Weilin Luo, Lei Sha, Zhifang Sui
AAAI2
2026 Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
abstract
Xin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, Zhenda Xie, Kezhao Huang, Xingkai Yu, Zhewen Hao, Han Zhang, Yu-Kun Li, Huishuai Zhang, Dongyan Zhao, Wenfeng Liang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xin Cheng 0002, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, Zhenda Xie, Kezhao Huang, Xingkai Yu, Zhewen Hao, Huishuai Zhang, Dongyan Zhao 0001, Wenfeng Liang
ACL (1)3
2025 Exploring Activation Patterns of Parameters in Language Models
abstract
Most work treats large language models as black boxes without an in-depth understanding of their internal working mechanism. To explain the internal representations of LLMs, we utilize a gradient-based metric to assess the activation level of model parameters. Based on this metric, we obtain three preliminary findings. (1) When the inputs are in the same domain, parameters in the shallow layers will be activated densely, which means a larger portion of parameters will have great impacts on the outputs. In contrast, parameters in the deep layers are activated sparsely. (2) When the inputs are across different domains, parameters in shallow layers exhibit higher similarity in the activation behavior than in deep layers. (3) In deep layers, the similarity of the distributions of activated parameters is positively correlated to the empirical data relevance. Further, we develop three validation experiments to solidify these findings. (1) Firstly, starting from the first finding, we attempt to configure different sparsities for different layers and find this method can benefit model pruning. (2) Secondly, we find that a pruned model based on one calibration set can better handle tasks related to the calibration task than those not related, which validates the second finding. (3) Thirdly, Based on the STS-B and SICK benchmarks, we find that two sentences with consistent semantics tend to share similar parameter activation patterns in deep layers, which aligns with our third finding. Our work sheds light on the behavior of parameter activation in LLMs, and we hope these findings will have the potential to inspire more practical applications.
Yudong Wang 0005, Damai Dai, Zhe Yang 0013, Jingyuan Ma, Zhifang Sui
AAAI2
2025 Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
abstract
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, Wangding Zeng. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo 0002, Liang Zhao 0026, Zhengyan Zhang, Zhenda Xie, Lean Wang, Zhiping Xiao 0001, Chong Ruan, Ming Zhang 0004, Wenfeng Liang, Wangding Zeng
ACL (1)3
2025 Language Models Encode the Value of Numbers Linearly
abstract
Large language models (LLMs) have exhibited impressive competence in various tasks, but their internal mechanisms on mathematical problems are still under-explored. In this paper, we study a fundamental question: how language models encode the value of numbers, a basic element in math. To study the question, we construct a synthetic dataset comprising addition problems and utilize linear probes to read out input numbers from the hidden states. Experimental results support the existence of encoded number values in LLMs on different layers, and these values can be extracted via linear probes. Further experiments show that LLMs store their calculation results in a similar manner, and we can intervene the output via simple vector additions, proving the causal connection between encoded numbers and language model outputs. Our research provides evidence that LLMs encode the value of numbers linearly, offering insights for better exploring, designing, and utilizing numeric information in LLMs.
Fangwei Zhu, Damai Dai, Zhifang Sui
COLING2
2025 Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
abstract
The rapid scaling of large language models (LLMs) has unveiled critical limitations in current hardware architectures, including constraints in memory capacity, computational efficiency, and interconnection bandwidth.DeepSeek-V3, trained on 2,048 NVIDIA H800 GPUs, demonstrates how hardware-aware model co-design can effectively address these challenges, enabling cost-efficient training and inference at scale.This paper presents an in-depth analysis of the DeepSeek-V3/R1 model architecture and its AI infrastructure, highlighting key innovations such as Multi-head Latent Attention (MLA) for enhanced memory efficiency, Mixture of Experts (MoE) architectures for optimized computation-communication trade-offs, FP8 mixed-precision training to unlock the full potential of hardware capabilities, and a Multi-Plane Network Topology to minimize * Yuqing Wang and Liyue Zhang are the corresponding authors of this paper.
Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Huazuo Gao, Jiashi Li, Panpan Huang, Shangyan Zhou, Shirong Ma, Wenfeng Liang, Ying He 0018, Yuxuan Liu 0019, Y. X. Wei
ISCA4
2024 DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
abstract
Damai Dai, Chengqi Deng, Chenggang Zhao, R.x. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y.k. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, Wenfeng Liang
ACL (1)1
2024 Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
abstract
Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Peiyi Wang, Lei Li 0039, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li 0005, Deli Chen, Zhifang Sui
ACL (1)5
2024 A Survey on In-context Learning
abstract
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, Zhifang Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Qingxiu Dong, Lei Li 0039, Damai Dai, Jingyuan Ma, Rui Li 0094, Heming Xia, Jingjing Xu 0001, Zhiyong Wu 0011, Baobao Chang, Xu Sun 0001, Lei Li 0005, Zhifang Sui
EMNLP3
2024 Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
abstract
Parameter-efficient fine-tuning (PEFT) is crucial for customizing Large Language Models (LLMs) with constrained resources.Although there have been various PEFT methods for dense-architecture LLMs, PEFT for sparsearchitecture LLMs is still underexplored.In this work, we study the PEFT method for LLMs with the Mixture-of-Experts (MoE) architecture and the contents of this work are mainly threefold: (1) We investigate the dispersion degree of the activated experts in customized tasks, and found that the routing distribution for a specific task tends to be highly concentrated, while the distribution of activated experts varies significantly across different tasks.(2) We propose Expert-Specialized Fine-Tuning, or ESFT, which tunes the experts most relevant to downstream tasks while freezing the other experts and modules; experimental results demonstrate that our method not only improves the tuning efficiency, but also matches or even surpasses the performance of fullparameter fine-tuning.(3) We further analyze the impact of the MoE architecture on expertspecialized fine-tuning.We find that MoE models with finer-grained experts are more advantageous in selecting the combination of experts that are most relevant to downstream tasks, thereby enhancing both the training efficiency and effectiveness.Our code is available at https://github.com/deepseek-ai/ESFT.
Zihan Wang 0010, Deli Chen, Damai Dai, Runxin Xu, Zhuoshu Li
EMNLP3
2023 Denoising Bottleneck with Mutual Information Maximization for Video Multimodal Fusion
abstract
Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu, Binghuai Lin, Yunbo Cao, Zhifang Sui. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui
ACL (1)2
2023 Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning
abstract
In-context learning (ICL) emerges as a promising capability of large language models (LLMs) by providing them with demonstration examples to perform diverse tasks.However, the underlying mechanism of how LLMs learn from the provided context remains under-explored.In this paper, we investigate the working mechanism of ICL through an information flow lens.Our findings reveal that label words in the demonstration examples function as anchors:(1) semantic information aggregates into label word representations during the shallow computation layers' processing; (2) the consolidated information in label words serves as a reference for LLMs' final predictions.Based on these insights, we introduce an anchor re-weighting method to improve ICL performance, a demonstration compression technique to expedite inference, and an analysis framework for diagnosing ICL errors in GPT2-XL.The promising applications of our findings again validate the uncovered ICL working mechanism and pave the way for future studies. 1
Lean Wang, Lei Li 0039, Damai Dai, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001
EMNLP3
2023 Neural Knowledge Bank for Pretrained Transformers
Damai Dai, Wenbin Jiang 0002, Qingxiu Dong, Yajuan Lyu, Zhifang Sui
NLPCC (2)1
2023 Mixture-of-Experts for Biomedical Question Answering
Damai Dai, Wenbin Jiang 0002, Yajuan Lyu, Zhifang Sui, Baobao Chang
NLPCC (1)1
2023 Coarse-to-Fine Entity Representations for Document-Level Relation Extraction
Damai Dai, Shuang Zeng, Baobao Chang, Zhifang Sui
NLPCC (2)1
2022 StableMoE: Stable Routing Strategy for Mixture of Experts
abstract
The Mixture-of-Experts (MoE) technique can scale up the model size of Transformers with an affordable computational overhead.We point out that existing learning-to-route MoE methods suffer from the routing fluctuation issue, i.e., the target expert of the same input may change along with training, but only one expert will be activated for the input during inference.The routing fluctuation tends to harm sample efficiency because the same input updates different experts but only one is finally used.In this paper, we propose STABLEMOE with two training stages to address the routing fluctuation problem.In the first training stage, we learn a balanced and cohesive routing strategy and distill it into a lightweight router decoupled from the backbone model.In the second training stage, we utilize the distilled router to determine the token-to-expert assignment and freeze it for a stable routing strategy.We validate our method on language modeling and multilingual machine translation.The results show that STABLEMOE outperforms existing MoE methods in terms of both convergence speed and performance.
Damai Dai, Li Dong 0004, Shuming Ma, Bo Zheng 0010, Zhifang Sui, Baobao Chang, Furu Wei
ACL (1)1
2022 Knowledge Neurons in Pretrained Transformers
abstract
Large-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus (Petroni et al., 2019; Jiang et al., 2020b).In this paper, we present preliminary studies on how factual knowledge is stored in pretrained Transformers by introducing the concept of knowledge neurons.Specifically, we examine the fill-in-the-blank cloze task for BERT.Given a relational fact, we propose a knowledge attribution method to identify the neurons that express the fact.We find that the activation of such knowledge neurons is positively correlated to the expression of their corresponding facts.In our case studies, we attempt to leverage knowledge neurons to edit (such as update, and erase) specific factual knowledge without fine-tuning.Our results shed light on understanding the storage of knowledge within pretrained Transformers.The code is available at https://github.com/ Hunter-DDM/knowledge-neurons.
Damai Dai, Li Dong 0004, Yaru Hao, Zhifang Sui, Baobao Chang, Furu Wei
ACL (1)1
2022 Robust Fine-tuning via Perturbation and Interpolation from In-batch Instances
abstract
Fine-tuning pretrained language models (PLMs) on downstream tasks has become common practice in natural language processing. However, most of the PLMs are vulnerable, e.g., they are brittle under adversarial attacks or imbalanced data, which hinders the application of the PLMs on some downstream tasks, especially in safe-critical scenarios. In this paper, we propose a simple yet effective fine-tuning method called Match-Tuning to force the PLMs to be more robust. For each instance in a batch, we involve other instances in the same batch to interact with it. To be specific, regarding the instances with other labels as a perturbation, Match-Tuning makes the model more robust to noise at the beginning of training. While nearing the end, Match-Tuning focuses more on performing an interpolation among the instances with the same label for better generalization. Extensive experiments on various tasks in GLUE benchmark show that Match-Tuning consistently outperforms the vanilla fine-tuning by 1.64 scores. Moreover, Match-Tuning exhibits remarkable robustness to adversarial attacks and data imbalance.
Shoujie Tong, Qingxiu Dong, Damai Dai, Yifan Song 0002, Tianyu Liu 0001, Baobao Chang, Zhifang Sui
IJCAI3
2022 On the Representation Collapse of Sparse Mixture of Experts
abstract
Sparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations. However, learning such a routing mechanism encourages token clustering around expert centroids, implying a trend toward representation collapse. In this work, we propose to estimate the routing scores between tokens and experts on a low-dimensional hypersphere. We conduct extensive experiments on cross-lingual language model pre-training and fine-tuning on downstream tasks. Experimental results across seven multilingual benchmarks show that our method achieves consistent gains. We also present a comprehensive analysis on the representation and routing behaviors of our models. Our method alleviates the representation collapse issue and achieves more consistent routing than the baseline mixture-of-experts methods.
Zewen Chi, Li Dong 0004, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xianling Mao, Heyan Huang, Furu Wei
NeurIPS4
2022 Plug-and-Play Module for Commonsense Reasoning in Machine Reading Comprehension
Damai Dai, Zhifang Sui, Baobao Chang
NLPCC (2)1
2021 Behind the Scenes: An Exploration of Trigger Biases Problem in Few-Shot Event Classification
abstract
Few-Shot Event Classification (FSEC) aims at developing a model for event prediction, which can generalize to new event types with a limited number of annotated data. Existing FSEC studies have achieved high accuracy on different benchmarks. However, we find they suffer from trigger biases that signify the statistical homogeneity between some trigger words and target event types, which we summarize as trigger overlapping and trigger separability. The biases can result in context-bypassing problem, i.e., correct classifications can be gained by looking at only the trigger words while ignoring the entire context. Therefore, existing models can be weak in generalizing to unseen data in real scenarios. To further uncover the trigger biases and assess the generalization ability of the models, we propose two new sampling methods, Trigger-Uniform Sampling (TUS) and COnfusion Sampling (COS), for the meta tasks construction during evaluation. Besides, to cope with the context-bypassing problem in FSEC models, we introduce adversarial training and trigger reconstruction techniques. Experiments show these techniques help not only improve the performance, but also enhance the generalization ability of models.
Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Damai Dai, Baobao Chang, Zhifang Sui
CIKM4
2021 Decompose, Fuse and Generate: A Formation-Informed Method for Chinese Definition Generation
abstract
Hua Zheng, Damai Dai, Lei Li, Tianyu Liu, Zhifang Sui, Baobao Chang, Yang Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Damai Dai, Lei Li 0039, Tianyu Liu 0001, Zhifang Sui, Baobao Chang, Yang Liu 0124
NAACL-HLT2
2019 LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts
abstract
We introduce the task of automatic live commenting. Live commenting, which is also called “video barrage”, is an emerging feature on online video sites that allows real-time comments from viewers to fly across the screen like bullets or roll at the right side of the screen. The live comments are a mixture of opinions for the video and the chit chats with other comments. Automatic live commenting requires AI agents to comprehend the videos and interact with human viewers who also make the comments, so it is a good testbed of an AI agent’s ability to deal with both dynamic vision and language. In this work, we construct a large-scale live comment dataset with 2,361 videos and 895,929 live comments. Then, we introduce two neural models to generate live comments based on the visual and textual contexts, which achieve better performance than previous neural baselines such as the sequence-to-sequence model. Finally, we provide a retrieval-based evaluation protocol for automatic live commenting where the model is asked to sort a set of candidate comments based on the log-likelihood score, and evaluated on metrics such as mean-reciprocal-rank. Putting it all together, we demonstrate the first “LiveBot”. The datasets and the codes can be found at https://github.com/lancopku/livebot.
Shuming Ma, Lei Cui 0001, Damai Dai, Furu Wei, Xu Sun 0001
AAAI3
2019 Learning to Control the Fine-grained Sentiment for Story Ending Generation
abstract
Automatic story ending generation is an interesting and challenging task in natural language generation.Previous studies are mainly limited to generate coherent, reasonable and diversified story endings, and few works focus on controlling the sentiment of story endings.This paper focuses on generating a story ending which meets the given fine-grained sentiment intensity.There are two major challenges to this task.First is the lack of story corpus which has fine-grained sentiment labels.Second is the difficulty of explicitly controlling sentiment intensity when generating endings.Therefore, we propose a generic and novel framework which consists of a sentiment analyzer and a sentimental generator, respectively addressing the two challenges.The sentiment analyzer adopts a series of methods to acquire sentiment intensities of the story dataset.The sentimental generator introduces the sentiment intensity into decoder via a Gaussian Kernel Layer to control the sentiment of the output.To the best of our knowledge, this is the first endeavor to control the fine-grained sentiment for story ending generation without manually annotating sentiment labels.Experiments show that our proposed framework can generate story endings which are not only more coherent and fluent but also able to meet the given sentiment intensity better. 1
Fuli Luo, Damai Dai, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Xu Sun 0001
ACL (1)2