EDBT 2026 Demo / reviewers in the wild / expert
Ning Ding 0002
dblp:04/4910-2
· DBLP profile ↗
49ranked-venue papers
8as first author
41since 2021 · last 2025
0000-0001-8758-9484ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 7 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fusing Highly Specialized Language Models for Comprehensive ExpertiseabstractNing Ding, Yulin Chen, Ganqu Cui, Xingtai Lv, Weilin Zhao, Kaiyan Zhang, Ruobing Xie, Bowen Zhou, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ning Ding 0002, Yulin Chen 0001, Ganqu Cui, Xingtai Lv, Weilin Zhao, Ruobing Xie, Bowen Zhou 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 1 |
| 2025 | Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single ProcessabstractSupervised Fine-Tuning (SFT) and Preference Optimization (PO) are key processes for aligning Language Models (LMs) with human preferences post pre-training.While SFT excels in efficiency and PO in effectiveness, they are often combined sequentially without integrating their optimization objectives.This approach ignores the opportunities to bridge their paradigm gap and take the strengths from both.In this paper, we interpret SFT and PO with two subprocesses -Preference Estimation and Transition Optimization -defined at token level within the Markov Decision Process (MDP).This modeling shows that SFT is only a special case of PO with inferior estimation and optimization.PO estimates the model's preference by its entire generation, while SFT only scores model's subsequent predicted tokens based on prior tokens from ground truth answer.These priors deviates from model's distribution, hindering the preference estimation and transition optimization.Building on this view, we introduce Intuitive Fine-Tuning (IFT) to integrate SFT and PO into a single process.Through a temporal residual connection, IFT brings better estimation and optimization by capturing LMs' intuitive sense of its entire answers.But it solely relies on a single policy and the same volume of non-preference-labeled data as SFT.Our experiments show that IFT performs comparably or even superiorly to SFT and some typical PO methods across several tasks, particularly those requires generation, reasoning, and fact-following abilities.An explainable Frozen Lake game further validates the effectiveness of IFT for getting competitive policy. Ermo Hua, Biqing Qi, Xingtai Lv, Ning Ding 0002, Bowen Zhou 0002 |
ACL (1) | 6 |
| 2025 | Advancing LLM Reasoning Generalists with Preference TreesabstractWe introduce EURUS, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B, Llama-3-8B, and Mixtral-8x22B, EURUS models achieve state-of-the-art results among open-source models on a diverse set of benchmarks covering mathematics, code generation, and logical reasoning problems. Notably, EURUX-8X22B outperforms GPT-3.5 Turbo in reasoning through a comprehensive benchmarking across 12 test sets covering five tasks. The strong performance of EURUS can be primarily attributed to ULTRAINTERACT, our newly-curated large-scale, high-quality training data dataset specifically designed for complex reasoning tasks. ULTRAINTERACT can be used in both supervised fine-tuning, preference learning, and reward modeling. It pairs each instruction with a preference tree consisting of (1) reasoning chains with diverse planning strategies in a unified format, (2) multi-turn interaction trajectories with the environment and the critique, and (3) pairwise positive and negative responses to facilitate preference learning. ULTRAINTERACT allows us to conduct an in-depth exploration of preference learning for reasoning tasks. Our investigation reveals that some well-established preference learning algorithms may be less suitable for reasoning tasks compared to their effectiveness in general conversations. The hypothesis is that in reasoning tasks, the space of correct answers is much smaller than that of incorrect ones, so it is necessary to explicitly increase the reward of chosen data. Therefore, in addition to increasing the reward margin as many preference learning algorithms do, the absolute values of positive responses’ rewards should be positive and may serve as a proxy for performance. Inspired by this, we derive a novel reward modeling objective and empirically that it leads to a stable reward modeling curve and better performance. Together with ULTRAINTERACT, we obtain a strong reward model. Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding 0002, Xingyao Wang 0002, Boji Shan, Zeyuan Liu, Ruobing Xie, Yankai Lin 0001, Zhenghao Liu 0001, Bowen Zhou 0002, Hao Peng 0015, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 4 |
| 2025 | OpenPRM: Building Open-domain Process-based Reward Models with Preference TreesabstractScaling inference-time computation is increasingly seen as the next frontier in scaling laws for large language models. Previous work in mathematics and coding has demonstrated the remarkable potential for inference-time scaling. During such scaling, fine-grained supervision through process-based reward models (PRMs) is essential for enhancement. However, exploration of inference-time scaling and PRMs in open-domain problems remains limited, where lacking exact answers and obtaining process supervision prove challenging. In this paper, we explore the construction of PRMs for open-domain tasks, specifically for instruction-following tasks. Utilizing existing outcome-based reward models (ORMs), we develop sentence-level preference trees based on the prefix similarity of parallel sampled candidates from datasets like UltraFeedback. This setup allows us to derive weak supervision for processes via back-propagation from outcome-level rewards. Subsequently, we integrate ORMs and PRMs under the same pairwise ranking objectives, resulting in our newly developed reward models, named OpenPRM. This approach significantly enhances the scalability of process-level supervision in open domains at minimal cost. We assess the performance of OpenPRM across various reward benchmarks, demonstrating its competitive edge over traditional ORMs in open domains and PRMs in specialized domains. Additionally, we investigate the scalability of inference-time computation for open-domain instructions. Our results highlight the limitations of ORMs’ scalability, while OpenPRM shows superior performance in scaled settings. Despite these advances, achieving automatic fine-grained supervision for open-domain inference-time scaling remains a substantial challenge. We hope these findings will spur further development of process supervision reward models in open-domain scenarios. Jiayuan Zhang 0001, Haoxin Li, Xuekai Zhu, Ermo Hua, Xingtai Lv, Ning Ding 0002, Biqing Qi, Bowen Zhou 0002 |
ICLR | 7 |
| 2025 | Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length GeneralizationabstractExtending the context length of Language Models (LMs) by improving Rotary Position Embedding (RoPE) has become a trend. While prior works mainly address RoPE’s limitations within attention, this paper uncovers the adverse effects on length generalization from nearly all parts of LMs. Using Discrete Signal Processing theory, we show that RoPE enables periodic attention by implicitly achieving Non-Uniform Discrete Fourier Transform. However, this periodicity is undermined by the spectrum damage caused by: 1) linear layers and activation functions outside of attention; 2) insufficiently trained frequency components brought by time-domain truncation. Building on our observations, we propose Fourier Position Embedding (FoPE), which enhances attention’s frequency-domain properties to improve both its periodic extension and length generalization. FoPE constructs Fourier Series and zero-outs the destructive frequency components, increasing model robustness against the spectrum damage. Experiments across various model scales and benchmarks show that, within varying context windows, FoPE maintains a more stable performance compared to other baselines. Several analyses and ablations bring further support to our method and theoretical modeling. Ermo Hua, Che Jiang, Xingtai Lv, Youbang Sun, Yuchen Fan 0001, Xuekai Zhu, Biqing Qi, Ning Ding 0002, Bowen Zhou 0002 |
ICML | 9 |
| 2025 | Free Process Rewards without Process LabelsabstractDifferent from its counterpart outcome reward models (ORMs), which evaluate the entire responses, a process reward model (PRM) scores a reasoning trajectory step by step, providing denser and more fine-grained rewards. However, training a PRM requires labels annotated at every intermediate step, presenting significant challenges for both manual and automatic data collection. This paper aims to address this challenge. Both theoretically and empirically, we show that an implicit PRM can be obtained at no additional cost, by simply training an ORM on the cheaper response-level labels. The only assumption is to parameterize the outcome reward as the log-likelihood ratios of the policy and reference models r$\phi$(y) = $\beta$ log $\pi$$\phi$(y) $\pi$ref(y) , which can be optimized regardless of the specific choice of loss objectives. In experiments, we instantiate our implicit PRMs with various objectives and evaluate their performance on MATH. We show that our implicit PRM outperforms a strong MCTS-based baseline á la Math-Shepherd (Wang et al., 2023) using less than 1/38 of the training data. Its performance can be further improved with majority voting. We further find that scaling up instructions and responses benefits our implicit PRM, and the latter brings a larger gain. Particularly, we find that our implicit PRM, when instantiated with the cross-entropy (CE) loss, is more data-efficient and can keep improving generation models even when trained with only one response per instruction, the setup that suffers from extreme data scarcity and imbalance. Further, instructions should be relevant to downstream tasks while the diversity of responses does not bring gains. Surprisingly, training on extra Math-Shepherd step labels brings no further improvements to our implicit PRM trained on only outcome data. We hope that our work will encourage a rethinking of PRM training approaches and contribute to making training PRMs more accessible. Lifan Yuan, Wendi Li, Huayu Chen, Ganqu Cui, Ning Ding 0002, Bowen Zhou 0002, Zhiyuan Liu 0001, Hao Peng 0001 |
ICML | 5 |
| 2025 | How to Synthesize Text Data without Model Collapse?abstractModel collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-$\{n\}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves data quality and enhances model performance. Xuekai Zhu, Daixuan Cheng, Hengli Li, Ermo Hua, Xingtai Lv, Ning Ding 0002, Zhouhan Lin, Zilong Zheng, Bowen Zhou 0002 |
ICML | 7 |
| 2025 | MedXpertQA: Benchmarking Expert-Level Medical Reasoning and UnderstandingabstractWe introduce MedXpertQA, a highly challenging and comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning. MedXpertQA includes 4,460 questions spanning 17 specialties and 11 body systems. It includes two subsets, Text for text evaluation and MM for multimodal evaluation. Notably, MM introduces expert-level exam questions with diverse images and rich clinical information, including patient records and examination results, setting it apart from traditional medical multimodal benchmarks with simple QA pairs generated from image captions. MedXpertQA applies rigorous filtering and augmentation to address the insufficient difficulty of existing benchmarks like MedQA, and incorporates specialty board questions to improve clinical relevance and comprehensiveness. We perform data synthesis to mitigate data leakage risk and conduct multiple rounds of expert reviews to ensure accuracy and reliability. We evaluate 18 leading models on MedXpertQA. Moreover, medicine is deeply connected to real-world decision-making, providing a rich and representative setting for assessing reasoning abilities beyond mathematics and code. To this end, we develop a reasoning-oriented subset to facilitate the assessment of o1-like models. Yuxin Zuo, Shang Qu, Zhang-Ren Chen, Xuekai Zhu, Ermo Hua, Ning Ding 0002, Bowen Zhou 0002 |
ICML | 8 |
| 2025 | The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE TrainingabstractRecent large language models (LLMs) exhibit impressive reasoning but often \textit{overthink}, generating excessively long responses that hinder efficiency. We introduce DIET (DIfficulty-AwarE Training), a framework that systematically cuts these "token calories" by integrating on-the-fly problem difficulty into the reinforcement learning (RL) process. DIET dynamically adapts token compression strategies by modulating token penalty strength and conditioning target lengths on estimated task difficulty, to optimize the performance-efficiency trade-off. We also theoretically analyze the pitfalls of naive reward weighting in group-normalized RL algorithms like GRPO, and propose \textit{Advantage Weighting} technique, which enables stable and effective implementation of these difficulty-aware objectives. Experimental results demonstrate that DIET significantly reduces token counts while simultaneously improving reasoning performance. Beyond raw token reduction, we show two crucial benefits largely overlooked by prior work: (1) DIET leads to superior \textbf{inference scaling}. By maintaining high per-sample quality with fewer tokens, it enables better scaling performance via majority voting under fixed computational budgets, an area where other methods falter. (2) DIET enhances the natural positive correlation between response length and problem difficulty, ensuring verbosity is appropriately allocated, unlike many existing compression methods that disrupt this relationship. Our analyses provide a principled and effective framework for developing more efficient, practical, and high-performing LLMs. Weize Chen, Jiarui Yuan, Tailin Jin, Ning Ding 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 4 |
| 2025 | Learning to Focus: Causal Attention Distillation via Gradient-Guided Token PruningabstractLarge language models (LLMs) have demonstrated significant improvements in contextual understanding. However, their ability to attend to truly critical information during long-context reasoning and generation still falls behind the pace. Specifically, our preliminary experiments reveal that certain distracting patterns can misdirect the model’s attention during inference, and removing these patterns substantially improves reasoning accuracy and generation quality. We attribute this phenomenon to spurious correlations in the training data, which obstruct the model’s capacity to infer authentic causal instruction–response relationships. This phenomenon may induce redundant reasoning processes, potentially resulting in significant inference overhead and, more critically, the generation of erroneous or suboptimal responses. To mitigate this, we introduce a two-stage framework called Learning to Focus (LeaF) leveraging intervention-based inference to disentangle confounding factors. In the first stage, LeaF employs gradient-based comparisons with an advanced teacher to automatically identify confounding tokens based on causal relationships in the training corpus. Then, in the second stage, it prunes these tokens during distillation to enact intervention, aligning the student’s attention with the teacher’s focus distribution on truly critical context tokens. Experimental results demonstrate that LeaF not only achieves an absolute improvement in various mathematical reasoning, code generation and multi-hop question answering benchmarks but also effectively suppresses attention to confounding tokens during inference, yielding a more interpretable and reliable reasoning model. Yiju Guo, Wenkai Yang, Zexu Sun, Ning Ding 0002, Zhiyuan Liu 0001, Yankai Lin 0001 |
NeurIPS | 4 |
| 2025 | DePass: Unified Feature Attributing by Simple Decomposed Forward PassabstractAttributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass. DePass decomposes hidden states into customized additive components, then propagates them with attention scores and MLP's activations fixed. It achieves faithful, fine-grained attribution without requiring auxiliary training. We validate DePass across token-level, model component-level, and subspace-level attribution tasks, demonstrating its effectiveness and fidelity. Our experiments highlight its potential to attribute information flow between arbitrary components of a Transformer model. We hope DePass serves as a foundational tool for broader applications in interpretability. Xiangyu Hong, Che Jiang, Biqing Qi, Youbang Sun, Ning Ding 0002, Bowen Zhou 0002 |
NeurIPS | 6 |
| 2025 | Scaling Physical Reasoning with the PHYSICS DatasetabstractLarge Language Models (LLMs) have achieved remarkable progress on advanced reasoning tasks such as mathematics and coding competitions. Meanwhile, physics, despite being both reasoning-intensive and essential to real-world understanding, received limited academic and industrial attention. This paper introduces PHYSICS, a dataset containing 16,568 high-quality physics problems spanning subjects and difficulty levels, to facilitate this issue. Specifically, PHYSICS is curated with exercises from over 100 textbooks through a carefully designed pipeline for quality control. It covers five major physics domains: Mechanics, Electromagnetism, Thermodynamics, Optics, and Modern Physics. It also spans a wide range of difficulty levels, from high school to graduate-level physics courses. To utilize the data for improving and evaluating the model's physical reasoning capabilities, we split the dataset into training and test sets, and provide reasoning paths generated by powerful reasoning models for the training data to facilitate model training. In addition, for the evaluation part, we find that existing evaluation frameworks exhibit biases in aspects such as units, simplification, and precision in physics domain. To balance efficiency and accuracy, we introduce a Rule+Model evaluation framework tailored to physics problems. Our evaluations on current state-of-the-art open-source and proprietary models highlight the limitations of current models in handling physics-related tasks. We hope that our dataset and evaluation methodology will jointly advance the development of LLMs in the field of physics. The code and data can be found at: https://github.com/Zhengsh123/PHYSICS. Shenghe Zheng, Qianjia Cheng, Junchi Yao, Mengsong Wu, Ning Ding 0002, Yu Cheng 0001, Shuyue Hu, Lei Bai 0001, Dongzhan Zhou, Ganqu Cui, Peng Ye 0006 |
NeurIPS | 6 |
| 2025 | TTRL: Test-Time Reinforcement LearningabstractThis paper investigates Reinforcement Learning (RL) on data without explicit labels for reasoning tasks in Large Language Models (LLMs). The core challenge of the problem is reward estimation during inference while not having access to ground-truth information. While this setting appears elusive, we find that common practices in Test-Time Scaling (TTS), such as majority voting, yield surprisingly effective rewards suitable for driving RL training. In this work, we introduce Test-Time Reinforcement Learning (TTRL), a novel method for training LLMs using RL on unlabeled data. TTRL enables self-evolution of LLMs by utilizing the priors in the pre-trained models. Our experiments demonstrate that TTRL consistently improves performance across a variety of tasks and models. Notably, TTRL boosts the pass@1 performance of Qwen-2.5-Math-7B by approximately 211% on the AIME 2024 with only unlabeled test data. Furthermore, although TTRL is only supervised by the Maj@N metric, TTRL has demonstrated performance to consistently surpass the upper limit of the initial model, and approach the performance of models trained directly on test data with ground-truth labels. Our experimental findings validate the general effectiveness of TTRL across various tasks and highlight TTRL's potential for broader tasks and domains. Yuxin Zuo, Shang Qu, Ganqu Cui, Xuekai Zhu, Haozhan Li, Xinwei Long, Ermo Hua, Biqing Qi, Youbang Sun, Zhiyuan Ma 0005, Lifan Yuan, Ning Ding 0002, Bowen Zhou 0002 |
NeurIPS | 15 |
| 2025 | Efficient Diffusion Models: A Comprehensive Survey From Principles to PracticesabstractAs one of the most popular and sought-after generative models in recent years, diffusion models have sparked the interests of many researchers and steadily shown excellent advantage in various generative tasks such as image synthesis, video generation, bioinformatics engineering, 3D scene rendering and multimodal generation, relying on their dense theoretical principles and reliable application practices. The remarkable success of these recent efforts on diffusion models comes largely from progressive design principles and efficient architecture, training, inference, and deployment methodologies. However, there has not been a comprehensive and in-depth review to summarize these principles and practices to help the rapid understanding and application of diffusion models. In this survey, we provide a new efficiency-oriented perspective on these existing efforts, which mainly focuses on the profound principles and efficient practices in architecture designs, model training, fast inference and reliable deployment, to guide further theoretical research, algorithm migration and model application for new scenarios in a reader-friendly way. Zhiyuan Ma 0005, Yuzhu Zhang, Guoli Jia, Yichao Ma, Gaofeng Liu, Ning Ding 0002, Jianjun Li 0010, Bowen Zhou 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning DatasetabstractHaoyu Wang, Shuo Wang, Yukun Yan, Xujia Wang, Zhiyu Yang, Yuzhuang Xu, Zhenghao Liu, Liner Yang, Ning Ding, Xu Han, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Shuo Wang 0013, Yukun Yan, Xujia Wang, Zhiyu Yang 0001, Yuzhuang Xu, Zhenghao Liu 0001, Liner Yang, Ning Ding 0002, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 9 |
| 2024 | CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction FollowingabstractWith the advancement of language models (LMs), their exposure to private data is increasingly inevitable, and their deployment (especially for smaller ones) on personal devices, such as PCs and smartphones, has become a prevailing trend.In contexts laden with user information, enabling models to both safeguard user privacy and execute commands efficiently emerges as an essential research imperative.In this paper, we propose CoGenesis, a collaborative generation framework integrating large (hosted on cloud infrastructure) and small models (deployed on local devices) to address privacy concerns logically.Initially, we design a pipeline to create personalized writing instruction datasets enriched with extensive context details as the testbed of this research issue.Subsequently, we introduce two variants of CoGenesis based on sketch and logits respectively.Our experimental findings, based on our synthesized dataset and two additional open-source datasets, indicate that: 1) Large-scale models perform well when provided with user context but struggle in the absence of such context.2) While specialized smaller models fine-tuned on the synthetic dataset show promise, they still lag behind their larger counterparts.3) Our CoGenesis framework, utilizing mixed-scale models, showcases competitive performance, providing a feasible solution to privacy issues.* Corresponding author 1 This paper defines large LMs (LLMs) as both closed and open-source models, designed for universal application and advanced performance, and intended for cloud deployment.Conversely, small LMs (SLMs) refer to models tailored for specific tasks and deployed on local devices. Jianyu Wang 0012, Ermo Hua, Biqing Qi, Ning Ding 0002, Bowen Zhou 0002 |
ACL (1) | 5 |
| 2024 | Empowering Private Tutoring by Chaining Large Language ModelsabstractArtificial intelligence has been applied in various aspects of online education to facilitate teaching and learning. However, few approaches have been made towards a complete AI-powered tutoring system. In this work, we explore the development of a full-fledged intelligent tutoring system based on large language models (LLMs). The proposed system ChatTutor, powered by state-of-the-art LLMs, is equipped with automatic course planning and adjusting, informative instruction, and adaptive quiz offering and evaluation. ChatTutor is decomposed into three inter-connected core processes: interaction, reflection, and reaction. Each process is implemented by chaining LLM-powered tools along with dynamically updated memory modules. To demonstrate the mechanism of each working module and the benefits of structured memory control and adaptive reflection, we conduct a wide range of analysis based on statistical results and user study. The analysis shows the designed processes boost system consistency and stability under long-term interaction and intentional disruptions, with up to 5% and 20% increase in performance respectively. Meanwhile, we also compare the system with scripts from real-world online learning platform and discuss the potential issues unique to LLM-based systems. Yulin Chen 0001, Ning Ding 0002, Hai-Tao Zheng 0002, Zhiyuan Liu 0001, Maosong Sun 0001, Bowen Zhou 0002 |
CIKM | 2 |
| 2024 | Controllable Preference Optimization: Toward Controllable Multi-Objective AlignmentabstractYiju Guo, Ganqu Cui, Lifan Yuan, Ning Ding, Zexu Sun, Bowen Sun, Huimin Chen, Ruobing Xie, Jie Zhou, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yiju Guo, Ganqu Cui, Lifan Yuan, Ning Ding 0002, Zexu Sun, Ruobing Xie, Jie Zhou 0016, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 4 |
| 2024 | Scalable Efficient Training of Large Language Models with Low-dimensional Projected AttentionabstractImproving the effectiveness and efficiency of large language models (LLMs) simultaneously is a critical yet challenging research goal.In this paper, we find that low-rank pre-training, normally considered as efficient methods that will compromise performance, can be scalably effective when reduced parameters are precisely targeted.Specifically, applying the low-dimensional module only to the attention layer -resolves this issue and enhances both effectiveness and efficiency.We refer to this structure as Low-dimensional Projected Attention (LPA) and provide an explanatory analysis.Through extensive experimentation at parameter scales of 130M, 370M, and scaling up to 3B, we have validated the effectiveness and scalability of LPA.Our results show that LPA model can save up to 12.4% in time while achieving an approximate 5% improvement in test perplexity (ppl) and on downstream tasks compared with the vanilla Transformer. Xingtai Lv, Ning Ding 0002, Ermo Hua, Ganqu Cui, Bowen Zhou 0002 |
EMNLP | 2 |
| 2024 | Predicting Emergent Abilities with Infinite Resolution EvaluationabstractThe scientific scale-up of large language models (LLMs) necessitates a comprehensive understanding of their scaling properties. However, the existing literature on the scaling properties only yields an incomplete answer: optimization loss decreases predictably as the model size increases, in line with established scaling law; yet no scaling law for task has been established and the task performances are far from predictable during scaling. Task performances typically show minor gains on small models until they improve dramatically once models exceed a size threshold, exemplifying the ''emergent abilities''. In this study, we discover that small models, although they exhibit minor performance, demonstrate critical and consistent task performance improvements that are not captured by conventional evaluation strategies due to insufficient measurement resolution. To measure such improvements, we introduce PassUntil, an evaluation strategy with theoretically infinite resolution, through massive sampling in the decoding phase. With PassUntil, we conduct a quantitative investigation into the scaling law of task performance. The investigation contains two parts. Firstly, a strict task scaling law that is not conventionally known to exist, is identified, enhancing the predictability of task performances. Remarkably, we are able to predict the performance of the 2.4B model on code generation with merely 0.05\% deviation before training starts, which is the first systematic attempt to verify predictable scaling proposed by GPT-4's report. Secondly, underpinned by PassUntil, we are able to study emergent abilities quantitatively. We identify a kind of accelerated emergence whose scaling curve cannot be fitted by standard scaling law function and has a increasing speed. We then examine two hypothesis and imply that the ``multiple circuits hypothesis'' might be responsible for the accelerated emergence. Shengding Hu, Xin Liu 0086, Xu Han 0007, Chaoqun He, Weilin Zhao, Yankai Lin 0001, Ning Ding 0002, Zebin Ou, Guoyang Zeng, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 8 |
| 2024 | KoLA: Carefully Benchmarking World Knowledge of Large Language ModelsabstractThe unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we believe meticulous and thoughtful designs are essential to thorough, unbiased, and applicable evaluations. Given the importance of world knowledge to LLMs, we construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) For ability modeling, we mimic human cognition to form a four-level taxonomy of knowledge-related abilities, covering 19 tasks. (2) For data, to ensure fair comparisons, we use both Wikipedia, a corpus prevalently pre-trained by LLMs, along with continuously collected emerging corpora, aiming to evaluate the capacity to handle unseen data and evolving knowledge. (3) For evaluation criteria, we adopt a contrastive system, including overall standard scores for better numerical comparability across tasks and models, and a unique self-contrast metric for automatically evaluating knowledge-creating ability. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings. The KoLA dataset will be updated every three months to provide timely references for developing LLMs and knowledge-related systems. Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Hao Peng 0015, Zijun Yao 0002, Hanming Li, Zheyuan Zhang 0002, Yushi Bai, Yantao Liu, Amy Xin, Kaifeng Yun, Linlu Gong, Nianyi Lin, Zhi-Li Wu, Yunjia Qi, Weikai Li 0002, Kaisheng Zeng, Ji Qi 0003, Hailong Jin, Jinxin Liu 0002, Yu Gu 0029, Yuan Yao 0011, Ning Ding 0002, Lei Hou 0001, Zhiyuan Liu 0001, Bin Xu 0001, Jie Tang 0001, Juan-Zi Li |
ICLR | 30 |
| 2024 | ULTRAFEEDBACK: Boosting Language Models with Scaled AI FeedbackabstractLearning from human feedback has become a pivot technique in aligning large language models (LLMs) with human preferences. However, acquiring vast and premium human feedback is bottlenecked by time, labor, and human capability, resulting in small sizes or limited topics of current datasets. This further hinders feedback learning as well as alignment research within the open-source community. To address this issue, we explore how to go beyond human feedback and collect high-quality AI feedback automatically for a scalable alternative. Specifically, we identify scale and diversity as the key factors for feedback data to take effect. Accordingly, we first broaden instructions and responses in both amount and breadth to encompass a wider range of user-assistant interactions. Then, we meticulously apply a series of techniques to mitigate annotation biases for more reliable AI feedback. We finally present UltraFeedback, a large-scale, high-quality, and diversified AI feedback dataset, which contains over 1 million GPT-4 feedback for 250k user-assistant conversations from various aspects. Built upon UltraFeedback, we align a LLaMA-based model by best-of-$n$ sampling and reinforcement learning, demonstrating its exceptional performance on chat benchmarks. Our work validates the effectiveness of scaled AI feedback data in constructing strong open-source chat language models, serving as a solid foundation for future feedback learning research. Ganqu Cui, Lifan Yuan, Ning Ding 0002, Guanming Yao, Bingxiang He, Wei Zhu 0016, Yuan Ni, Guo Tong Xie, Ruobing Xie, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICML | 3 |
| 2024 | UltraMedical: Building Specialized Generalists in BiomedicineabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities across various domains and are moving towards more specialized areas. Recent advanced proprietary models such as GPT-4 and Gemini have achieved significant advancements in biomedicine, which have also raised privacy and security challenges. The construction of specialized generalists hinges largely on high-quality datasets, enhanced by techniques like supervised fine-tuning and reinforcement learning from human or AI feedback, and direct preference optimization. However, these leading technologies (e.g., preference learning) are still significantly limited in the open source community due to the scarcity of specialized data. In this paper, we present the UltraMedical collections, which consist of high-quality manual and synthetic datasets in the biomedicine domain, featuring preference annotations across multiple advanced LLMs. By utilizing these datasets, we fine-tune a suite of specialized medical models based on Llama-3 series, demonstrating breathtaking capabilities across various medical benchmarks. Moreover, we develop powerful reward models skilled in biomedical and general reward benchmark, enhancing further online preference learning within the biomedical LLM community. Sihang Zeng, Ermo Hua, Ning Ding 0002, Zhang-Ren Chen, Zhiyuan Ma 0005, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, Xingtai Lv, Jinfang Hu, Zhiyuan Liu 0001, Bowen Zhou 0002 |
NeurIPS | 4 |
| 2024 | Estimation of control area in badminton doubles with pose information from top and back view drone videosabstractAbstract The application of visual tracking to the performance analysis of sports players in dynamic competitions is vital for effective coaching. In doubles matches, coordinated positioning is crucial for maintaining control of the court and minimizing opponents’ scoring opportunities. The analysis of such teamwork plays a vital role in understanding the dynamics of the game. However, previous studies have primarily focused on analyzing and assessing singles players without considering occlusion in broadcast videos. These studies have relied on discrete representations, which involve the analysis and representation of specific actions (e.g., strokes) or events that occur during the game while overlooking the meaningful spatial distribution. In this work, we present the first annotated drone dataset from top and back views in badminton doubles and propose a framework to estimate the control area probability map, which can be used to evaluate teamwork performance. We present an efficient framework of deep neural networks that enables the calculation of full probability surfaces. This framework utilizes the embedding of a Gaussian mixture map of players’ positions and employs graph convolution on their poses. In the experiment, we verify our approach by comparing various baselines and discovering the correlations between the score and control area. Additionally, we propose a practical application for assessing optimal positioning to provide instructions during a game. Our approach offers both visual and quantitative evaluations of players’ movements, thereby providing valuable insights into doubles teamwork. The dataset and related project code is available at https://github.com/Ning-D/Drone_BD_ControlArea Ning Ding 0002, Kazuya Takeda, Wenhui Jin, Yingjiu Bei, Keisuke Fujii 0001 |
Multim. Tools Appl. | 1 |
| 2024 | Exploring Universal Intrinsic Task Subspace for Few-Shot Learning via Prompt TuningabstractWhy can pre-trained language models (PLMs) learn universal representations and effectively adapt to broad NLP tasks differing a lot superficially? In this work, we empirically find evidence indicating that the adaptations of PLMs to various few-shot tasks can be reparameterized as optimizing only a few free parameters in a unified low-dimensionalintrinsic task subspace, which may help us understand why PLMs could easily adapt to various NLP tasks with small-scale data. To find such a subspace and examine its universality, we propose an analysis pipeline calledintrinsic prompt tuning(IPT). Specifically, we resort to the recent success of prompt tuning and decompose the soft prompts of multiple NLP tasks into the same low-dimensional nonlinear subspace, then we learn to adapt the PLM to unseen data or tasks by only tuning parameters in this subspace. In the experiments, we study diverse few-shot NLP tasks and surprisingly find that in a 250-dimensional subspace found with 100 tasks, by only tuning 250 free parameters, we can recover 97% and 83% of the full prompt tuning performance for 100 seen tasks (using different training data) and 20 unseen tasks, respectively, showing great generalization ability of the found intrinsic task subspace. Besides being an analysis tool, IPTcould further help us improve the prompt tuning stability. Yujia Qin, Xiaozhi Wang, Yusheng Su, Yankai Lin 0001, Ning Ding 0002, Jing Yi, Weize Chen, Zhiyuan Liu 0001, Juan-Zi Li, Lei Hou 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Exploring Lottery Prompts for Pre-trained Language ModelsabstractYulin Chen, Ning Ding, Xiaobin Wang, Shengding Hu, Haitao Zheng, Zhiyuan Liu, Pengjun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yulin Chen 0001, Ning Ding 0002, Xiaobin Wang, Shengding Hu, Hai-Tao Zheng 0002, Zhiyuan Liu 0001, Pengjun Xie |
ACL (1) | 2 |
| 2023 | Decoder Tuning: Efficient Language Understanding as DecodingabstractWith the evergrowing sizes of pre-trained models (PTMs), it has been an emerging practice to only provide the inference APIs for users, namely model-as-a-service (MaaS) setting.To adapt PTMs with model parameters frozen, most current approaches focus on the input side, seeking for powerful prompts to stimulate models for correct answers.However, we argue that input-side adaptation could be arduous due to the lack of gradient signals and they usually require thousands of API queries, resulting in high computation and time costs.In light of this, we present Decoder Tuning (DecT), which in contrast optimizes task-specific decoder networks on the output side.Specifically, DecT first extracts prompt-stimulated output scores for initial predictions.On top of that, we train an additional decoder network on the output representations to incorporate posterior data knowledge.By gradientbased optimization, DecT can be trained within several seconds and requires only one PTM query per sample.Empirically, we conduct extensive natural language understanding experiments and show that DecT significantly outperforms state-of-the-art algorithms with a 200× speed-up.Our codes are available at https://github.com/thunlp/DecT. Ganqu Cui, Ning Ding 0002, Longtao Huang, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 3 |
| 2023 | WebCPM: Interactive Web Search for Chinese Long-form Question AnsweringabstractYujia Qin, Zihan Cai, Dian Jin, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin, Xu Han, Ning Ding, Huadong Wang, Ruobing Xie, Fanchao Qi, Zhiyuan Liu, Maosong Sun, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yujia Qin, Zihan Cai, Lan Yan, Shihao Liang, Kunlun Zhu, Yankai Lin 0001, Xu Han 0007, Ning Ding 0002, Ruobing Xie, Fanchao Qi, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0024 |
ACL (1) | 9 |
| 2023 | Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsabstractFine-tuning on instruction data has been widely validated as an effective practice for implementing chat language models like ChatGPT.Scaling the diversity and quality of such data, although straightforward, stands a great chance of leading to improved performance.This paper aims to push the upper bound of opensource models further.We first provide a systematically designed, diverse, informative, large-scale dataset of instructional conversations, UltraChat, which does not involve human queries.Our objective is to capture the breadth of interactions between a human user and an AI assistant and employs a comprehensive framework to generate multi-turn conversation iteratively.UltraChat contains 1.5 million high-quality multi-turn dialogues and covers a wide range of topics and instructions.Our statistical analysis of UltraChat reveals its superiority in various key metrics, including scale, average length, diversity, coherence, etc., solidifying its position as a leading opensource dataset.Building upon UltraChat, we fine-tune a LLaMA model to create a powerful conversational model, UltraLM.Our evaluations indicate that UltraLM consistently outperforms other open-source models, including WizardLM and Vicuna, the previously recognized state-of-the-art open-source models. Ning Ding 0002, Yulin Chen 0001, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu 0001, Maosong Sun 0001, Bowen Zhou 0002 |
EMNLP | 1 |
| 2023 | Sparse Low-rank Adaptation of Pre-trained Language ModelsabstractFine-tuning pre-trained large language models in a parameter-efficient manner is widely studied for its effectiveness and efficiency.The popular method of low-rank adaptation (LoRA) offers a notable approach, hypothesizing that the adaptation process is intrinsically low-dimensional.Although LoRA has demonstrated commendable performance, it is implemented with a fixed and unalterable intrinsic rank that might not always be the ideal choice.Recognizing the need for more flexible adaptation, we extend the methodology of LoRA to an innovative approach we call sparse low-rank adaptation (SoRA) that enables dynamic adjustments to the intrinsic rank during the adaptation process.We achieve this through the incorporation of a gate unit optimized with proximal gradient method in the training stage, controlling the cardinality of rank under the sparsity of the gate.In the subsequent inference stage, we eliminate the parameter blocks corresponding to the zeroed-out ranks, to reduce each SoRA module back to a concise yet rankoptimal LoRA.Our approach strengthens the representation power of LoRA by initializing it with a higher rank, while efficiently taming a temporarily increased number of parameters via updating in a sparse way.We further introduce a sparsifying scheduler for SoRA, aiming to examine the impact of the number of nonzero parameters on the model's memorization and generalization.Our experimental results demonstrate that SoRA can outperform other baselines even with 70% retained parameters and 70% training time. Ning Ding 0002, Xingtai Lv, Qiaosen Wang, Yulin Chen 0001, Bowen Zhou 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 1 |
| 2023 | Exploring the Impact of Model Scaling on Parameter-Efficient TuningabstractYusheng Su, Chi-Min Chan, Jiali Cheng, Yujia Qin, Yankai Lin, Shengding Hu, Zonghan Yang, Ning Ding, Xingzhi Sun, Guotong Xie, Zhiyuan Liu, Maosong Sun. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yusheng Su, Chi-Min Chan, Jiali Cheng, Yujia Qin, Yankai Lin 0001, Shengding Hu, Zonghan Yang, Ning Ding 0002, Xingzhi Sun 0002, Guo Tong Xie, Zhiyuan Liu 0001, Maosong Sun 0001 |
EMNLP | 8 |
| 2023 | CRaSh: Clustering, Removing, and Sharing Enhance Fine-tuning without Full Large Language ModelabstractInstruction tuning has recently been recognized as an effective way of aligning Large Language Models (LLMs) to enhance their generalization ability across various tasks.However, when tuning publicly accessible, centralized LLMs with private instruction data, privacy concerns are inevitable.While direct transfer of parameterized modules between models is a plausible approach to address this, its implications and effectiveness need further exploration.This paper focuses on Offsite-Tuning (OFT), a representative technique that transfers transformer blocks between centralized LLMs and downstream emulators.Given the limited understanding of the underlying mechanism of OFT, we perform an empirical analysis on LLMs from the perspectives of representation and functional similarity.Interestingly, our findings reveal a unique modular structure within the layers of LLMs that appears to emerge as the model size expands.Simultaneously, we note subtle but potentially significant changes in representation and intermediate predictions across the layers.Inspired by these observations, we propose CRaSh, involving Clustering, Removing, and Sharing, a training-free strategy to derive improved emulators from LLMs.CRaSh significantly boosts performance of OFT with billions of parameters.Furthermore, we investigate the optimal solutions yielded by fine-tuning with and without full model through the lens of loss landscape.Our findings demonstrate a linear connectivity among these optima falling over the same basin, thereby highlighting the effectiveness of CRaSh and OFT.The source code is publicly available at https://github.com/TsinghuaC3I/CRaSh. Ning Ding 0002, Biqing Qi, Xuekai Zhu, Xinwei Long, Bowen Zhou 0002 |
EMNLP | 2 |
| 2022 | Prototypical Verbalizer for Prompt-based Few-shot TuningabstractPrompt-based tuning for pre-trained language models (PLMs) has shown its effectiveness in few-shot learning.Typically, prompt-based tuning wraps the input text into a cloze question.To make predictions, the model maps the output words to labels via a verbalizer, which is either manually designed or automatically built.However, manual verbalizers heavily depend on domain-specific prior knowledge and human efforts, while finding appropriate label words automatically still remains challenging.In this work, we propose the prototypical verbalizer (ProtoVerb) which is built directly from training data.Specifically, Pro-toVerb learns prototype vectors as verbalizers by contrastive learning.In this way, the prototypes summarize training instances and are able to enclose rich class-level semantics.We conduct experiments on both topic classification and entity typing tasks, and the results demonstrate that ProtoVerb significantly outperforms current automatic verbalizers, especially when training data is extremely scarce.More surprisingly, ProtoVerb consistently boosts promptbased tuning even on untuned PLMs, indicating an elegant non-tuning way to utilize PLMs.Our codes are avaliable at https: //github.com/thunlp/OpenPrompt. Ganqu Cui, Shengding Hu, Ning Ding 0002, Longtao Huang, Zhiyuan Liu 0001 |
ACL (1) | 3 |
| 2022 | Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text ClassificationabstractShengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu, Jingang Wang, Juanzi Li, Wei Wu, Maosong Sun. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shengding Hu, Ning Ding 0002, Zhiyuan Liu 0001, Jingang Wang, Juan-Zi Li, Wei Wu 0014, Maosong Sun 0001 |
ACL (1) | 2 |
| 2022 | MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation ExtractionabstractXiaozhi Wang, Yulin Chen, Ning Ding, Hao Peng, Zimu Wang, Yankai Lin, Xu Han, Lei Hou, Juanzi Li, Zhiyuan Liu, Peng Li, Jie Zhou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Xiaozhi Wang, Yulin Chen 0001, Ning Ding 0002, Hao Peng 0015, Yankai Lin 0001, Xu Han 0007, Lei Hou 0001, Juan-Zi Li, Zhiyuan Liu 0001, Peng Li 0030, Jie Zhou 0016 |
EMNLP | 3 |
| 2022 | ProQA: Structural Prompt-based Pre-training for Unified Question AnsweringabstractWanjun Zhong, Yifan Gao, Ning Ding, Yujia Qin, Zhiyuan Liu, Ming Zhou, Jiahai Wang, Jian Yin, Nan Duan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Wanjun Zhong, Yifan Gao 0001, Ning Ding 0002, Yujia Qin, Zhiyuan Liu 0001, Ming Zhou 0001, Jiahai Wang, Jian Yin 0001, Nan Duan 0001 |
NAACL-HLT | 3 |
| 2022 | Sparse Structure Search for Delta TuningabstractAdapting large pre-trained models (PTMs) through fine-tuning imposes prohibitive computational and storage burdens. Recent studies of delta tuning (DT), i.e., parameter-efficient tuning, find that only optimizing a small portion of parameters conditioned on PTMs could yield on-par performance compared to conventional fine-tuning. Generally, DT methods exquisitely design delta modules (DT modules) which could be applied to arbitrary fine-grained positions inside PTMs. However, the effectiveness of these fine-grained positions largely relies on sophisticated manual designation, thereby usually producing sub-optimal results. In contrast to the manual designation, we explore constructing DT modules in an automatic manner. We automatically \textbf{S}earch for the \textbf{S}parse \textbf{S}tructure of \textbf{Delta} Tuning (S$^3$Delta). Based on a unified framework of various DT methods, S$^3$Delta conducts the differentiable DT structure search through bi-level optimization and proposes shifted global sigmoid method to explicitly control the number of trainable parameters. Extensive experiments show that S$^3$Delta surpasses manual and random structures with less trainable parameters. The searched structures preserve more than 99\% fine-tuning performance with 0.01\% trainable parameters. Moreover, the advantage of S$^3$Delta is amplified with extremely low trainable parameters budgets (0.0009\%$\sim$0.01\%). The searched structures are transferable and explainable, providing suggestions and guidance for the future design of DT methods. Our codes are publicly available at \url{https://github.com/thunlp/S3Delta}. Shengding Hu, Zhen Zhang 0008, Ning Ding 0002, Yadao Wang, Yasheng Wang, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 3 |
| 2021 | Few-NERD: A Few-shot Named Entity Recognition DatasetabstractNing Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, Zhiyuan Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ning Ding 0002, Yulin Chen 0001, Xiaobin Wang, Xu Han 0007, Pengjun Xie, Hai-Tao Zheng 0002, Zhiyuan Liu 0001 |
ACL/IJCNLP (1) | 1 |
| 2021 | CLINE: Contrastive Learning with Semantic Negative Examples for Natural Language UnderstandingabstractDong Wang, Ning Ding, Piji Li, Haitao Zheng. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ning Ding 0002, Piji Li, Hai-Tao Zheng 0002 |
ACL/IJCNLP (1) | 2 |
| 2021 | Prototypical Representation Learning for Relation Extraction
Ning Ding 0002, Xiaobin Wang, Rui Wang 0005, Pengjun Xie, Ying Shen 0001, Fei Huang 0002, Hai-Tao Zheng 0002, Rui Zhang 0003 |
ICLR | 1 |
| 2021 | Modeling Relation Paths for Knowledge Graph CompletionabstractKnowledge graphs (KG) often encounter knowledge incompleteness. The path reasoning that predicts the unknown path relation between pairwise entities based on existing facts is one of the most promising approaches to the knowledge graph completion. However, most conventional path reasoning methods exclusively consider the entity description included in fact triples, ignoring both the type information of entities and the interaction between different semantic representations. In this study, we propose a novel method, Type-aware Attentive Path Reasoning (TAPR), to complete the knowledge graph by simultaneously considering KG structural information, textual information, and type information. More specifically, we first leverage types to enrich the representational learning of entities and relationships. Next, we describe a type-level attention to select the most relevant type of given entity in a specific triple without any predefined rules or patterns to reduce the impact of noisy types. After learning the distributed representation of all paths, path-level attention assigns different weights to paths, from which relations among entity pairs are calculated. We conduct a series of experiments on a real-world dataset to demonstrate the effectiveness of TAPR. Experimental results show that our method significantly outperforms all baselines on link prediction and entity prediction tasks. Ying Shen 0001, Ning Ding 0002, Hai-Tao Zheng 0002, Yaliang Li, Min Yang 0007 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Integrating Linguistic Knowledge to Sentence Paraphrase GenerationabstractParaphrase generation aims to rewrite a text with different words while keeping the same meaning. Previous work performs the task based solely on the given dataset while ignoring the availability of external linguistic knowledge. However, it is intuitive that a model can generate more expressive and diverse paraphrase with the help of such knowledge. To fill this gap, we propose Knowledge-Enhanced Paraphrase Network (KEPN), a transformer-based framework that can leverage external linguistic knowledge to facilitate paraphrase generation. (1) The model integrates synonym information from the external linguistic knowledge into the paraphrase generator, which is used to guide the decision on whether to generate a new word or replace it with a synonym. (2) To locate the synonym pairs more accurately, we adopt an incremental encoding scheme to incorporate position information of each synonym. Besides, a multi-task architecture is designed to help the framework jointly learn the selection of synonym pairs and the generation of expressive paraphrase. Experimental results on both English and Chinese datasets show that our method significantly outperforms the state-of-the-art approaches in terms of both automatic and human evaluation. Zibo Lin, Ziran Li, Ning Ding 0002, Hai-Tao Zheng 0002, Ying Shen 0001, Wei Wang 0138, Cong-Zhi Zhao |
AAAI | 3 |
| 2020 | Coupling Distant Annotation and Adversarial Training for Cross-Domain Chinese Word SegmentationabstractFully supervised neural approaches have achieved significant progress in the task of Chinese word segmentation (CWS).Nevertheless, the performance of supervised models tends to drop dramatically when they are applied to outof-domain data.Performance degradation is caused by the distribution gap across domains and the out of vocabulary (OOV) problem.In order to simultaneously alleviate these two issues, this paper proposes to couple distant annotation and adversarial training for crossdomain CWS.For distant annotation, we rethink the essence of "Chinese words" and design an automatic distant annotation mechanism that does not need any supervision or pre-defined dictionaries from the target domain.The approach could effectively explore domain-specific words and distantly annotate the raw texts for the target domain.For adversarial training, we develop a sentence-level training procedure to perform noise reduction and maximum utilization of the source domain information.Experiments on multiple realworld datasets across various domains show the superiority and robustness of our model, significantly outperforming previous state-ofthe-art cross-domain CWS methods. Ning Ding 0002, Dingkun Long, Muhua Zhu, Pengjun Xie, Xiaobin Wang, Hai-Tao Zheng 0002 |
ACL | 1 |
| 2020 | Hierarchy-Aware Global Model for Hierarchical Text ClassificationabstractHierarchical text classification is an essential yet challenging subtask of multi-label text classification with a taxonomic hierarchy.Existing methods have difficulties in modeling the hierarchical label structure in a global view.Furthermore, they cannot make full use of the mutual interactions between the text feature space and the label space.In this paper, we formulate the hierarchy as a directed graph and introduce hierarchy-aware structure encoders for modeling label dependencies.Based on the hierarchy encoder, we propose a novel end-to-end hierarchy-aware global model (Hi-AGM) with two variants.A multi-label attention variant (HiAGM-LA) learns hierarchyaware label embeddings through the hierarchy encoder and conducts inductive fusion of labelaware text features.A text feature propagation model (HiAGM-TP) is proposed as the deductive variant that directly feeds text features into hierarchy encoders.Compared with previous works, both HiAGM-LA and HiAGM-TP achieve significant and consistent improvements on three benchmark datasets. Jie Zhou 0013, Chunping Ma, Dingkun Long, Ning Ding 0002, Pengjun Xie, Gongshen Liu |
ACL | 5 |
| 2020 | Infobox-to-text Generation with Tree-like Planning based Attention NetworkabstractWe study the problem of infobox-to-text generation that aims to generate a textual description from a key-value table. Representing the input infobox as a sequence, previous neural methods using end-to-end models without order-planning suffer from the problems of incoherence and inadaptability to disordered input. Recent planning-based models only implement static order-planning to guide the generation, which may cause error propagation between planning and generation. To address these issues, we propose a Tree-like PLanning based Attention Network (Tree-PLAN) which leverages both static order-planning and dynamic tuning to guide the generation. A novel tree-like tuning encoder is designed to dynamically tune the static order-plan for better planning by merging the most relevant attributes together layer by layer. Experiments conducted on two datasets show that our model outperforms previous methods on both automatic and human evaluation, and demonstrate that our model has better adaptability to disordered input. Ziran Li, Ning Ding 0002, Ying Shen 0001, Hai-Tao Zheng 0002 |
IJCAI | 3 |
| 2020 | Triple-to-Text Generation with an Anchor-to-Prototype FrameworkabstractGenerating a textual description from a set of RDF triplets is a challenging task in natural language generation. Recent neural methods have become the mainstream for this task, which often generate sentences from scratch. However, due to the huge gap between the structured input and the unstructured output, the input triples alone are insufficient to decide an expressive and specific description. In this paper, we propose a novel anchor-to-prototype framework to bridge the gap between structured RDF triples and natural text. The model retrieves a set of prototype descriptions from the training data and extracts writing patterns from them to guide the generation process. Furthermore, to make a more precise use of the retrieved prototypes, we employ a triple anchor that aligns the input triples into groups so as to better match the prototypes. Experimental results on both English and Chinese datasets show that our method significantly outperforms the state-of-the-art baselines in terms of both automatic and manual evaluation, demonstrating the benefit of learning guidance from retrieved prototypes to facilitate triple-to-text generation. Ziran Li, Zibo Lin, Ning Ding 0002, Hai-Tao Zheng 0002, Ying Shen 0001 |
IJCAI | 3 |
| 2020 | Generalized Local Aggregation for Large Scale Gaussian Process RegressionabstractDespite being one of the most popular nonparametric approaches, Gaussian process regression (GPR) suffers from O(n3) computational burden and the computation is infeasible for large-scale scenarios. To reduce the computational complexity, many Shannon-mutual-information-based aggregation methods were proposed, whereas these methods can not effectively identify the importance of experts in some cases. To address this problem, we generalize the traditional mutual information-based methods (GPoE, RBCM, GRBCM) based on Tsallis mutual information. Accordingly, the generated weight distribution is more sparse tending to focus on those experts with good performance. To obtain adaptive and data-dependent entropic-index in Tsallis entropy, we propose three heuristic algorithms to solve our model. Extensive experiments show that, the proposed method can improve the prediction of both the mean and variance, and the improvement of variance prediction is significant in many cases. Yinghua Gao, Naiqi Li, Ning Ding 0002, Yiming Li 0004, Tao Dai 0001, Shutao Xia |
IJCNN | 3 |
| 2019 | Chinese Relation Extraction with Multi-Grained Information and External Linguistic KnowledgeabstractChinese relation extraction is conducted using neural networks with either character-based or word-based inputs, and most existing methods typically suffer from segmentation errors and ambiguity of polysemy.To address the issues, we propose a multi-grained lattice framework (MG lattice) for Chinese relation extraction to take advantage of multi-grained language information and external linguistic knowledge.In this framework, (1) we incorporate word-level information into character sequence inputs so that segmentation errors can be avoided.(2) We also model multiple senses of polysemous words with the help of external linguistic knowledge, so as to alleviate polysemy ambiguity.Experiments on three realworld datasets in distinct domains show consistent and significant superiority and robustness of our model, as compared with other baselines.The source code of this paper can be obtained from https://github.com/ thunlp/Chinese_NRE. Ziran Li, Ning Ding 0002, Zhiyuan Liu 0001, Hai-Tao Zheng 0002, Ying Shen 0001 |
ACL (1) | 2 |
| 2019 | Event Detection with Trigger-Aware Lattice Neural NetworkabstractNing Ding, Ziran Li, Zhiyuan Liu, Haitao Zheng, Zibo Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ning Ding 0002, Ziran Li, Zhiyuan Liu 0001, Hai-Tao Zheng 0002, Zibo Lin |
EMNLP/IJCNLP (1) | 1 |