Yeyun Gong

dblp:06/10400 · DBLP profile ↗
← Back
80ranked-venue papers
5as first author
50since 2021 · last 2026
0000-0001-9954-9674ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 75 · 4 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
abstract
Multilingual Alignment is an effective and representative paradigm to enhance LLMs' multilingual capabilities, which transfers the capabilities from the high-resource languages to the low-resource languages. Meanwhile, some research on language-specific neurons provides a new perspective to analyze and understand LLMs' mechanisms. However, we find that there are many neurons that are shared by multiple but not all languages and cannot be correctly classified. In this work, we propose a ternary classification methodology that categorizes neurons into three types, including language-specific neurons, language-related neurons, and general neurons. And we propose a corresponding identification algorithm to distinguish these different types of neurons. Furthermore, based on the distributional characteristics of different types of neurons, we divide the LLMs' internal process for multilingual inference into four parts: (1) multilingual understanding, (2) shared semantic space reasoning, (3) multilingual output space transformation, and (4) vocabulary space outputting. Additionally, we systematically analyze the models before and after alignment with a focus on different types of neurons. We also analyze the phenomenon of ''Spontaneous Multilingual Alignment''. Overall, our work conducts a comprehensive investigation based on different types of neurons, providing empirical results and valuable insights to better understand multilingual alignment and multilingual capabilities of LLMs.
Shimao Zhang, Zhejian Lai, Xiang Liu 0023, Shuaijie She, Yeyun Gong, Shujian Huang, Jiajun Chen 0001
AAAI6
2026 Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions
abstract
Boyan Duan, Xiao Liang, Shuai Lu, Yaoxiang Wang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Mao Yang, Weizhu Chen, Yeyun Gong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Boyan Duan, Yaoxiang Wang, Yelong Shen, Kai-Wei Chang 0001, Ying Nian Wu, Mao Yang 0004, Weizhu Chen, Yeyun Gong
ACL (1)10
2026 Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
abstract
Zhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji, Peijie Wang, Haotian Xu, Xing W, Haizhen Huang, Weiwei Deng, Yeyun Gong, Zhijiang Guo, Xiao Liu, Fei Yin, Cheng-Lin Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhongzhi Li, Lei Ji 0001, Xing W, Haizhen Huang, Yeyun Gong, Zhijiang Guo, Xiao Liu 0029, Cheng-Lin Liu 0001
ACL (1)10
2026 Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability
abstract
Xiao Liang, Zhong-Zhi Li, Zhenghao Lin, Eric Hanchen Jiang, Hengyuan Zhang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Yeyun Gong, Weizhu Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhongzhi Li, Zhenghao Lin, Eric Hanchen Jiang, Yelong Shen, Kai-Wei Chang 0001, Ying Nian Wu, Yeyun Gong, Weizhu Chen
ACL (1)9
2026 Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
abstract
Kailai Yang, Xiao Liu, Lei Ji, Hao Li, Xiao Liang, Zhiwei Liu, Yeyun Gong, Peng Cheng, Mao Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Kailai Yang, Xiao Liu 0029, Lei Ji 0001, Hao Li 0074, Zhiwei Liu 0003, Yeyun Gong, Peng Cheng 0005, Mao Yang 0004
ACL (1)7
2025 Key-Point-Driven Data Synthesis with Its Enhancement on Mathematical Reasoning
abstract
Large language models have shown great potential in complex reasoning tasks, yet their performance is often hampered by the scarcity of high-quality and reasoning-focused training datasets. Addressing this challenge, we propose Key-PointDriven Data Synthesis (KPDDS), a novel data synthesis framework that synthesizes question-answer pairs by leveraging key points and exemplar practices from authentic data sources. KPDDS ensures the generation of novel questions with rigorous quality control and substantial scalability. As a result, we present KPMath, an extensive synthetic dataset tailored for mathematical reasoning, comprising over 800K questionanswer pairs. Utilizing KPMath and augmenting it with additional reasoning-intensive corpora, we create the comprehensive KPMath-Plus dataset. Our experiments demonstrate that this dataset can enhance the mathematical reasoning performance of models across various architectures and sizes. The Qwen1.5-72B model, fine-tuned on KPMath-Plus, achieves 87.0% accuracy on GSM8K and 58.3% on MATH, surpassing competitors in the 7B to 72B range and best commercial models like GPT-4 across multiple math reasoning datasets.
Xiao Liu 0029, Yeyun Gong, Zhibin Gou, Yelong Shen, Nan Duan 0001, Weizhu Chen
AAAI3
2025 Enhancing Large Language Model Performance with Gradient-Based Parameter Selection
abstract
Large language models (LLMs) have revolutionized numerous fields of research, driving significant advancements in natural language processing, machine translation, and beyond. Although the extensive number of parameters contributes a lot to the great success, existing studies indicate that not all model parameters hold equal importance, which further leads to redundancy during the parameter update process. Recent works for reducing redundant parameter updates for LLMs either lack task-specific data information, may leading to suboptimal model performance, or discard transformer components or insignificant parameters, limiting the model's scalability across different tasks and potentially compromising the LLM structure. To address these issues and further enhance the performance of LLMs, we propose Gradient-Mask Tuning (GMT), a method that selectively updates parameters based on gradient information, which is specific to the target tasks. Specifically, after calculating gradients during back propagation, we measure their absolute values and mask those with small absolute values. Our empirical results in various training paradigms like SFT and DPO for various domains of tasks demonstrate that GMT not only preserves the original network structure but also enhances the potential performance of LLMs. Further analysis indicates that GMT exhibits insensitivity to mask ratio and possesses computational efficiency comparable to vanilla training approach.
Haoling Li, Yeyun Gong
AAAI4
2025 Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
abstract
It is well-known that a diverse corpus is critical for training large language models, which are typically constructed from a mixture of various domains.In general, previous efforts resort to either sampling training data from different domains with static proportions or dynamically adjusting these proportions during training to optimise pretraining performance.However, few methods addressed the complexity of domain-adaptive continual pre-training.To fill this gap, we propose Velocitune, a novel framework that dynamically assesses learning velocity and adjusts data proportions accordingly, favouring slower learning domains while de-emphasising faster learning ones, which is guided by a scaling law to estimate the desired learning goal for each domain with a less associated cost.To evaluate the effectiveness of Velocitune, we conduct experiments on a dataset focused on reasoning tasks with CodeLlama, as well as on a corpus of system commands using Llama3 and Mistral.Velocitune achieves performance gains in both math and code reasoning tasks and command-line generation benchmarks.Further analysis reveals that key factors driving the effectiveness of Velocitune include target estimation and data ordering.* The first two authors contribute equally.
Zheheng Luo, Haoling Li, Yeyun Gong
ACL (1)5
2025 Automated Proof Generation for Rust Code via Self-Evolution
abstract
Ensuring correctness is crucial for code generation. Formal verification offers a definitive assurance of correctness, but demands substantial human effort in proof construction and hence raises a pressing need for automation. The primary obsta- cle lies in the severe lack of data—there is much fewer proofs than code snippets for Large Language Models (LLMs) to train upon. In this paper, we introduce SAFE, a framework that overcomes the lack of human-written proofs to enable automated proof generation of Rust code. SAFE establishes a self-evolving cycle where data synthesis and fine-tuning collaborate to enhance the model capability, leveraging the definitive power of a symbolic verifier in telling correct proofs from incorrect ones. SAFE also re-purposes the large number of synthesized incorrect proofs to train the self-debugging capability of the fine-tuned models, empowering them to fix incorrect proofs based on the verifier’s feedback. SAFE demonstrates superior efficiency and precision compared to GPT-4o. Through tens of thousands of synthesized proofs and the self-debugging mechanism, we improve the capa- bility of open-source models, initially unacquainted with formal verification, to automatically write proofs for Rust code. This advancement leads to a signifi- cant improvement in performance, achieving a 52.52% accuracy rate in a bench- mark crafted by human experts, a significant leap over GPT-4o’s performance of 14.39%.
Tianyu Chen 0006, Shan Lu 0001, Yeyun Gong, Chenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Hao Yu 0016, Nan Duan 0001, Peng Cheng 0005, Fan Yang 0024, Shuvendu K. Lahiri, Tao Xie 0001, Lidong Zhou
ICLR4
2025 Integrative Decoding: Improving Factuality via Implicit Self-consistency
abstract
Self-consistency-based approaches, which involve repeatedly sampling multiple outputs and selecting the most consistent one as the final response, prove to be remarkably effective in improving the factual accuracy of large language models. Nonetheless, existing methods usually have strict constraints on the task format, largely limiting their applicability. In this paper, we present Integrative Decoding (ID), to unlock the potential of self-consistency in open-ended generation tasks. ID operates by constructing a set of inputs, each prepended with a previously sampled response, and then processes them concurrently, with the next token being selected by aggregating of all their corresponding predictions at each decoding step. In essence, this simple approach implicitly incorporates self-consistency in the decoding objective. Extensive evaluation shows that ID consistently enhances factuality over a wide range of language models, with substantial improvements on the TruthfulQA (+11.2%), Biographies (+15.4%) and LongFact (+8.5%) benchmarks. The performance gains amplify progressively as the number of sampled responses increases, indicating the potential of ID to scale up with repeated sampling.
Yeyun Gong, Yuji Zhang 0002, Kaishuai Xu, Wenge Liu, Wenjie Li 0002, Jian Jiao 0007, Qi Chen 0009, Peng Cheng 0005, Wayne Xiong
ICLR3
2025 Alchemy: Amplifying Theorem-Proving Capability Through Symbolic Mutation
abstract
Formal proofs are challenging to write even for experienced experts. Recent progress in Neural Theorem Proving (NTP) shows promise in expediting this process. However, the formal corpora available on the Internet are limited compared to the general text, posing a significant data scarcity challenge for NTP. To address this issue, this work proposes Alchemy, a general framework for data synthesis that constructs formal theorems through symbolic mutation. Specifically, for each candidate theorem in Mathlib, we identify all invocable theorems that can be used to rewrite or apply to it. Subsequently, we mutate the candidate theorem by replacing the corresponding term in the statement with its equivalent form or antecedent. As a result, our method increases the number of theorems in Mathlib by an order of magnitude, from 110k to 6M. Furthermore, we perform continual pretraining and supervised finetuning on this augmented corpus for large language models. Experimental results demonstrate the effectiveness of our approach, achieving a 4.70% absolute performance improvement on Leandojo benchmark. Additionally, our approach achieves a 2.47% absolute performance gain on the out-of-distribution miniF2F benchmark based on the synthetic data. To provide further insights, we conduct a comprehensive analysis of synthetic data composition and the training paradigm, offering valuable guidance for developing a strong theorem prover.
Shaonan Wu, Yeyun Gong, Nan Duan 0001, Ping Wei 0001
ICLR3
2025 Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
abstract
Using large teacher models to guide the training of smaller student models has become the prevailing paradigm for efficient and effective learning. However, vocabulary mismatches between teacher and student language models pose significant challenges in language modeling, resulting in divergent token sequences and output distributions. To overcome these limitations, we propose Vocabulary-agnostic Teacher Guided Language Modeling (VocAgnoLM), a novel approach that bridges the gap caused by vocabulary mismatch through two key methods: (1) Token-level Lexical Alignment, which aligns token sequences across mismatched vocabularies, and (2) Teacher Guided Loss, which leverages the loss of teacher model to guide effective student training. We demonstrate its effectiveness in language modeling with 1B student model using various 7B teacher models with different vocabularies. Notably, with Qwen2.5-Math-Instruct, a teacher model sharing only about 6% of its vocabulary with TinyLlama, VocAgnoLM achieves a 46% performance improvement compared to naive continual pretraining. Furthermore, we demonstrate that VocAgnoLM consistently benefits from stronger teacher models, providing a robust solution to vocabulary mismatches in language modeling.
Haebin Shin, Lei Ji 0001, Xiao Liu 0029, Yeyun Gong
ICML4
2025 Optimizing Large Language Model Training Using FP4 Quantization
abstract
The growing computational demands of training large language models (LLMs) necessitate more efficient methods. Quantized training presents a promising solution by enabling low-bit arithmetic operations to reduce these costs. While FP8 precision has demonstrated feasibility, leveraging FP4 remains a challenge due to significant quantization errors and limited representational capacity. This work introduces the first FP4 training framework for LLMs, addressing these challenges with two key innovations: a differentiable quantization estimator for precise weight updates and an outlier clamping and compensation strategy to prevent activation collapse. To ensure stability, the framework integrates a mixed-precision training scheme and vector-wise quantization. Experimental results demonstrate that our FP4 framework achieves accuracy comparable to BF16 and FP8, with minimal degradation, scaling effectively to 13B-parameter LLMs trained on up to 100B tokens. With the emergence of next-generation hardware supporting FP4, our framework sets a foundation for efficient ultra-low precision training.
Yeyun Gong, Xiao Liu 0029, Guoshuai Zhao 0001, Ziyue Yang 0002, Baining Guo, Zhengjun Zha, Peng Cheng 0005
ICML2
2025 Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
abstract
Sungjin Park, Xiao Liu, Yeyun Gong, Edward Choi. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Xiao Liu 0029, Yeyun Gong, Edward Choi 0003
NAACL (Long Papers)3
2025 Generative Prompt Internalization
abstract
Haebin Shin, Lei Ji, Yeyun Gong, Sungdong Kim, Eunbi Choi, Minjoon Seo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Haebin Shin, Lei Ji 0001, Yeyun Gong, Sungdong Kim, Eunbi Choi, Minjoon Seo
NAACL (Long Papers)3
2025 SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for training large language models (LLMs) on complex reasoning tasks, such as mathematical problem solving. A prerequisite for the scalability of RLVR is a high-quality problem set with precise and verifiable answers. However, the scarcity of well-crafted human-labeled math problems and limited-verification answers in existing distillation-oriented synthetic datasets limit their effectiveness in RL. Additionally, most problem synthesis strategies indiscriminately expand the problem set without considering the model’s capabilities, leading to low efficiency in generating useful questions. To mitigate this issue, we introduce a Self-aware Weakness-driven problem Synthesis framework (SwS) that systematically identifies model deficiencies and leverages them for problem augmentation. Specifically, we define weaknesses as questions that the model consistently fails to learn through its iterative sampling during RL training. We then extract the core concepts from these failure cases and synthesize new problems to strengthen the model's weak areas in subsequent augmented training, enabling it to focus on and gradually overcome its weaknesses. Without relying on external knowledge distillation, our framework enables robust generalization by empowering the model to self-identify and address its weaknesses in RL, yielding average performance gains of 10% and 7.7% on 7B and 32B models across eight mainstream reasoning benchmarks. Our code and data are available at https://anonymous.4open.science/r/SwS-E6F5/
Zhongzhi Li, Yeyun Gong, Yelong Shen, Ying Nian Wu, Weizhu Chen
NeurIPS3
2025 Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
abstract
State Space Models (SSMs) offer remarkable performance gains in efficient sequence modeling, with constant per-step inference-time computation and memory complexity. Recent advances, such as Mamba, further enhance SSMs with input-dependent gating and hardware-aware implementations, positioning them as strong alternatives to Transformers for long sequence modeling. However, efficiently scaling the expressive power of SSMs, particularly with Mixture of Experts (MoE), remains challenging, as naive integration attempts often falter or degrade performance. In this work, we introduce Routing Mamba (RoM), a novel approach that scales SSM parameters using sparse mixtures of linear projection experts. By sharing routing decisions between projection layers and lightweight sub-modules within Mamba across experts, RoM leverages synergies among linear projection experts for effective and efficient sparse scaling of Mamba layers. At a scale of 1.3B active parameters (10B total) and 16K training sequence length, RoM achieves language modeling performance equivalent to a dense Mamba model requiring over 2.3$\times$ more active parameters, and demonstrates consistent perplexity across context lengths. Experimental results further show RoM effectively scales hybrid language models, yielding a 23% FLOPS saving compared to dense Mamba scaling for similar performance. We release our training codebase at https://github.com/zhanzheng8585/Routing-Mamba.
Zheng Zhan 0001, Liliang Ren, Shuohang Wang, Yeyun Gong, Yanzhi Wang 0001, Yelong Shen
NeurIPS6
2025 PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
abstract
Inspired by the impressive reasoning capabilities demonstrated by reinforcement learning approaches like DeepSeek-R1, recent emerging research has begun exploring the use of reinforcement learning (RL) to enhance vision-language models (VLMs) for multimodal reasoning tasks. However, most existing multimodal reinforcement learning approaches remain limited to spatial reasoning within single-image contexts, yet still struggle to generalize to more complex and real-world scenarios involving multi-image positional reasoning, where understanding the relationships across images is crucial. To address this challenge, we propose a general reinforcement learning approach PeRL tailored for interleaved multimodal tasks, and a multi-stage strategy designed to enhance the exploration-exploitation trade-off, thereby improving learning efficiency and task performance. Specifically, we introduce permutation of image sequences to simulate varied positional relationships to explore more spatial and positional diversity. Furthermore, we design a rollout filtering mechanism for resampling to focus on trajectories that contribute most to learning optimal behaviors to exploit learned policies effectively. We evaluate our model on 5 widely-used multi-image benchmarks and 3 single-image benchmarks. Our experiments confirm that PeRL trained model consistently surpasses R1-related and interleaved VLM baselines by a large margin, achieving state-of-the-art performance on multi-image benchmarks, while preserving comparable performance on single-image tasks.
Shuoshuo Zhang, Haoling Li, Zhongzhi Li, Jie Wu 0001, Lei Ji 0001, Yeyun Gong, Yelong Shen, Yujiu Yang 0001
NeurIPS10
2025 From intention to implementation: automating biomedical research via LLMs
Linghang Shi, Aobo Zhuang, Yeyun Gong
Sci. China Inf. Sci.5
2025 AutoVerus: Automated Proof Generation for Rust Code
abstract
Generative AI has shown its value for many software engineering tasks. Still in its infancy, large language model (LLM)-based proof generation lags behind LLM-based code generation. In this paper, we present A uto V erus . A uto V erus uses LLMs to automatically generate correctness proof for Rust code. A uto V erus is designed to match the unique features of Verus, a verification tool that can prove the correctness of Rust code using proofs and specifications also written in Rust. A uto V erus consists of a network of agents that are crafted and orchestrated to mimic human experts’ three phases of proof construction: preliminary proof generation, proof refinement guided by generic tips, and proof debugging guided by verification errors. To thoroughly evaluate A uto V erus and help foster future research in this direction, we have built a benchmark suite of 150 non-trivial proof tasks, based on existing code-generation benchmarks and verification benchmarks. Our evaluation shows that A uto V erus can automatically generate correct proof for more than 90% of them, with more than half of them tackled in less than 30 seconds or 3 LLM calls.
Chenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Jianan Yao, Weidong Cui, Yeyun Gong, Chris Hawblitzel, Shuvendu K. Lahiri, Jacob R. Lorch, Fan Yang 0024, Ziqiao Zhou, Shan Lu 0001
Proc. ACM Program. Lang.6
2024 PROM: A Phrase-level Copying Mechanism with Pre-training for Abstractive Summarization
abstract
Based on the remarkable achievements of pre-trained language models in abstractive summarization, the copying mechanism has proved helpful by improving the factuality, stability, and overall performance. This work proposes PROM, a new PhRase-level cOpying Mechanism that enhances attention on n-grams, which can be applied to zero-shot summarization with pre-training. PROM adds an indicator layer to explicitly pick up tokens in n-gram that can be copied from the source, and calculates an auxiliary loss for the copying prediction. Empirical studies show that PROM makes significant improvements in fine-tuning on benchmarks. In the zero-shot setting, PROM is utilized in the self-supervised pre-training on raw corpora and provides new general baselines on a wide range of summarization datasets. Further analysis shows that PROM performs more reasonable copying and contributes to faithfulness. Our code is publicly available at https://github.com/xbmxb/PROM.
Xinbei Ma, Yeyun Gong, Hai Zhao 0001, Nan Duan 0001
LREC/COLING2
2024 APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning
abstract
Long-form numerical reasoning aims to generate a reasoning program to calculate the answer for a given question. Previous work followed a retriever-generator framework, where the retriever selects key facts from a long-form document, and the generator generates a reasoning program based on the retrieved facts. However, they treated all facts equally without considering the different contributions of facts with and without numerical information. Furthermore, they ignored program consistency, leading to the wrong punishment of programs that differed from the ground truth. In order to address these issues, we proposed APOLLO (An optimized training aPproach fOr Long-form numericaL reasOning), to improve long-form numerical reasoning. APOLLO includes a number-aware negative sampling strategy for the retriever to discriminate key numerical facts, and a consistency-based reinforcement learning with target program augmentation for the generator to ultimately increase the execution accuracy. Experimental results on the FinQA and ConvFinQA leaderboards verify the effectiveness of our proposed methods, achieving the new state-of-the-art.
Jiashuo Sun, Hang Zhang 0029, Chen Lin 0001, Xiangdong Su, Yeyun Gong, Jian Guo 0016
LREC/COLING5
2024 Knowledge Enhanced Pre-training for Cross-lingual Dense Retrieval
abstract
In recent years, multilingual pre-trained language models (mPLMs) have achieved significant progress in cross-lingual dense retrieval. However, most mPLMs neglect the importance of knowledge. Knowledge always conveys similar semantic concepts in a language-agnostic manner, while query-passage pairs in cross-lingual retrieval also share common factual information. Motivated by this observation, we introduce KEPT, a novel mPLM that effectively leverages knowledge to learn language-agnostic semantic representations. To achieve this, we construct a multilingual knowledge base using hyperlinks and cross-language page alignment data annotated by Wiki. From this knowledge base, we mine intra- and cross-language pairs by extracting symmetrically linked segments and multilingual entity descriptions. Subsequently, we adopt contrastive learning with the mined pairs to pre-train KEPT. We evaluate KEPT on three widely-used benchmarks, considering both zero-shot cross-lingual transfer and supervised multilingual fine-tuning scenarios. Extensive experimental results demonstrate that KEPT achieves strong multilingual and cross-lingual retrieval performance with significant improvements over existing mPLMs.
Hang Zhang 0029, Yeyun Gong, Dayiheng Liu, Shunyu Zhang, Xingwei He 0003, Jiancheng Lv 0001, Jian Guo 0016
LREC/COLING2
2024 Task Oriented In-Domain Data Augmentation
abstract
Large Language Models (LLMs) have shown superior performance in various applications and fields.To achieve better performance on specialized domains such as law and advertisement, LLMs are often continue pre-trained on in-domain data.However, existing approaches suffer from two major issues.First, in-domain data are scarce compared with general domainagnostic data.Second, data used for continual pre-training are not task-aware, such that they may not be helpful to downstream applications.We propose TRAIT, a task-oriented in-domain data augmentation framework.Our framework is divided into two parts: in-domain data selection and task-oriented synthetic passage generation.The data selection strategy identifies and selects a large amount of in-domain data from general corpora, and thus significantly enriches domain knowledge in the continual pre-training data.The synthetic passages contain guidance on how to use domain knowledge to answer questions about downstream tasks.By training on such passages, the model aligns with the need of downstream applications.We adapt LLMs to two domains: advertisement and math.On average, TRAIT improves LLM performance by 8% in the advertisement domain and 7.5% in the math domain.
Simiao Zuo, Yeyun Gong, Qiang Lou, Yi Liu 0071, Shao-Lun Huang, Jian Jiao 0007
EMNLP4
2024 CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
abstract
Recent developments in large language models (LLMs) have been impressive. However, these models sometimes show inconsistencies and problematic behavior, such as hallucinating facts, generating flawed code, or creating offensive and toxic content. Unlike these models, humans typically utilize external tools to cross-check and refine their initial content, like using a search engine for fact-checking, or a code interpreter for debugging. Inspired by this observation, we introduce a framework called CRITIC that allows LLMs, which are essentially “black boxes” to validate and progressively amend their own outputs in a manner similar to human interaction with tools. More specifically, starting with an initial output, CRITIC interacts with appropriate tools to evaluate certain aspects of the text, and then revises the output based on the feedback obtained during this validation process. Comprehensive evaluations involving free-form question answering, mathematical program synthesis, and toxicity reduction demonstrate that CRITIC consistently enhances the performance of LLMs. Meanwhile, our research highlights the crucial importance of external feedback in promoting the ongoing self-improvement of LLMs.
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang 0001, Nan Duan 0001, Weizhu Chen
ICLR3
2024 ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
abstract
Large language models have made significant progress in various language tasks, yet they still struggle with complex mathematics. In this paper, we propose ToRA a series of Tool-integrated Reasoning Agents designed to solve challenging mathematical problems by seamlessly integrating natural language reasoning with the utilization of external tools (e.g., computation libraries and symbolic solvers), thereby amalgamating the analytical prowess of language and the computational efficiency of tools. To train ToRA, we curate interactive tool-use trajectories on mathematical datasets, apply imitation learning on the annotations, and propose output space shaping to further refine models' reasoning behavior. As a result, ToRA models significantly outperform open-source models on 10 mathematical reasoning datasets across all scales with 13%-19% absolute improvements on average. Notably, ToRA-7B reaches 44.6% on the competition-level dataset MATH, surpassing the best open-source model WizardMath-70B by 22% absolute. ToRA-34B is also the first open-source model that achieves an accuracy exceeding 50% on MATH, which significantly outperforms GPT-4's CoT result, and is competitive with GPT-4 solving problems with programs. Additionally, we conduct a comprehensive analysis of the benefits and remaining challenges of tool interaction for mathematical reasoning, providing valuable insights for future research.
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang 0001, Minlie Huang, Nan Duan 0001, Weizhu Chen
ICLR3
2024 Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph
abstract
Although large language models (LLMs) have achieved significant success in various tasks, they often struggle with hallucination problems, especially in scenarios requiring deep and responsible reasoning. These issues could be partially addressed by introducing external knowledge graphs (KG) in LLM reasoning. In this paper, we propose a new LLM-KG integrating paradigm ``$\hbox{LLM}\otimes\hbox{KG}$'' which treats the LLM as an agent to interactively explore related entities and relations on KGs and perform reasoning based on the retrieved knowledge. We further implement this paradigm by introducing a new approach called Think-on-Graph (ToG), in which the LLM agent iteratively executes beam search on KG, discovers the most promising reasoning paths, and returns the most likely reasoning results. We use a number of well-designed experiments to examine and illustrate the following advantages of ToG: 1) compared with LLMs, ToG has better deep reasoning power; 2) ToG has the ability of knowledge traceability and knowledge correctability by leveraging LLMs reasoning and expert feedback; 3) ToG provides a flexible plug-and-play framework for different LLMs, KGs and prompting strategies without any additional training cost; 4) the performance of ToG with small LLM models could exceed large LLM such as GPT-4 in certain scenarios and this reduces the cost of LLM deployment and application. As a training-free method with lower computational cost and better generality, ToG achieves overall SOTA in 6 out of 9 datasets where most previous SOTAs rely on additional training.
Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin 0001, Yeyun Gong, Lionel M. Ni, Harry Shum, Jian Guo 0016
ICLR6
2024 Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models
abstract
Yi Luo, Zhenghao Lin, YuHao Zhang, Jiashuo Sun, Chen Lin, Chengjin Xu, Xiangdong Su, Yelong Shen, Jian Guo, Yeyun Gong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhenghao Lin, Jiashuo Sun, Chen Lin 0001, Chengjin Xu, Xiangdong Su, Yelong Shen, Jian Guo 0016, Yeyun Gong
NAACL-HLT10
2024 Not All Tokens Are What You Need for Pretraining
abstract
Previous language model pre-training methods have uniformly applied a next-token prediction loss to all training tokens. Challenging this norm, we posit that ''Not all tokens in a corpus are equally important for language model training''. Our initial analysis examines token-level training dynamics of language model, revealing distinct loss patterns for different tokens. Leveraging these insights, we introduce a new language model called Rho-1. Unlike traditional LMs that learn to predict every next token in a corpus, Rho-1 employs Selective Language Modeling (SLM), which selectively trains on useful tokens that aligned with the desired distribution. This approach involves scoring training tokens using a reference model, and then training the language model with a focused loss on tokens with higher scores. When continual continual pretraining on 15B OpenWebMath corpus, Rho-1 yields an absolute improvement in few-shot accuracy of up to 30% in 9 math tasks. After fine-tuning, Rho-1-1B and 7B achieved state-of-the-art results of 40.6% and 51.8% on MATH dataset, respectively - matching DeepSeekMath with only 3% of the pretraining tokens. Furthermore, when continual pretraining on 80B general tokens, Rho-1 achieves 6.8% average enhancement across 15 diverse tasks, increasing both data efficiency and performance of the language model pre-training.
Zhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu 0029, Yelong Shen, Ruochen Xu, Chen Lin 0001, Yujiu Yang 0001, Jian Jiao 0007, Nan Duan 0001, Weizhu Chen
NeurIPS3
2024 LEAD: Liberal Feature-based Distillation for Dense Retrieval
abstract
Knowledge distillation is often used to transfer knowledge from a strong teacher model to a relatively weak student model. Traditional methods include response-based methods and feature-based methods. Response-based methods are widely used but suffer from lower upper limits of performance due to their ignorance of intermediate signals, while feature-based methods have constraints on vocabularies, tokenizers and model architectures. In this paper, we propose a liberal feature-based distillation method (LEAD). LEAD aligns the distribution between the intermediate layers of teacher model and student model, which is effective, extendable, portable and has no requirements on vocabularies, tokenizers, or model architectures. Extensive experiments show the effectiveness of LEAD on widely-used benchmarks, including MS MARCO Passage Ranking, TREC 2019 DL Track, MS MARCO Document Ranking and TREC 2020 DL Track. Our code is available in https://github.com/microsoft/SimXNS/tree/main/LEAD.
Hao Sun 0015, Xiao Liu 0029, Yeyun Gong, Anlei Dong, Jingwen Lu, Yan Zhang 0117, Linjun Yang, Rangan Majumder, Nan Duan 0001
WSDM3
2023 CAPSTONE: Curriculum Sampling for Dense Retrieval with Document Expansion
abstract
The dual-encoder has become the de facto architecture for dense retrieval.Typically, it computes the latent representations of the query and document independently, thus failing to fully capture the interactions between the query and document.To alleviate this, recent research has focused on obtaining query-informed document representations.During training, it expands the document with a real query, but during inference, it replaces the real query with a generated one.This inconsistency between training and inference causes the dense retrieval model to prioritize query information while disregarding the document when computing the document representation.Consequently, it performs even worse than the vanilla dense retrieval model because its performance heavily relies on the relevance between the generated queries and the real query.In this paper, we propose a curriculum sampling strategy that utilizes pseudo queries during training and progressively enhances the relevance between the generated query and the real query.By doing so, the retrieval model learns to extend its attention from the document alone to both the document and query, resulting in high-quality queryinformed document representations.Experimental results on both in-domain and out-ofdomain datasets demonstrate that our approach outperforms previous dense retrieval models.
Xingwei He 0003, Yeyun Gong, A-Long Jin, Hang Zhang 0029, Anlei Dong, Jian Jiao 0007, Siu-Ming Yiu, Nan Duan 0001
EMNLP2
2023 Query Rewriting in Retrieval-Augmented Large Language Models
abstract
Large Language Models (LLMs) play powerful, black-box readers in the retrieve-thenread pipeline, making remarkable progress in knowledge-intensive tasks.This work introduces a new framework, Rewrite-Retrieve-Read instead of the previous retrieve-then-read for the retrieval-augmented LLMs from the perspective of the query rewriting.Unlike prior studies focusing on adapting either the retriever or the reader, our approach pays attention to the adaptation of the search query itself, for there is inevitably a gap between the input text and the needed knowledge in retrieval.We first prompt an LLM to generate the query, then use a web search engine to retrieve contexts.Furthermore, to better align the query to the frozen modules, we propose a trainable scheme for our pipeline.A small language model is adopted as a trainable rewriter to cater to the black-box LLM reader.The rewriter is trained using the feedback of the LLM reader by reinforcement learning.Evaluation is conducted on downstream tasks, open-domain QA and multiple-choice QA.Experiments results show consistent performance improvement, indicating that our framework is proven effective and scalable, and brings a new framework for retrieval-augmented LLM 1 .
Xinbei Ma, Yeyun Gong, Hai Zhao 0001, Nan Duan 0001
EMNLP2
2023 Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph Denoise
abstract
In this paper, we introduce a novel dIffusion language modEl pre-training framework for text generation, which we call GENIE. GENIE is a large-scale pre-trained diffusion language model that consists of an encoder and a diffusion-based decoder, which can generate text by gradually transforming a random noise sequence into a coherent text sequence. To pre-train GENIE on a large-scale language corpus, we design a new continuous paragraph denoise objective, which encourages the diffusion-decoder to reconstruct a clean text paragraph from a corrupted version, while preserving the semantic and syntactic coherence. We evaluate GENIE on four downstream text generation benchmarks, namely XSum, CNN/DailyMail, Gigaword, and CommonGen. Our experimental results show that GENIE achieves comparable performance with the state-of-the-art autoregressive models on these benchmarks, and generates more diverse text samples. The code and models of GENIE are available at https://github.com/microsoft/ProphetNet/tree/master/GENIE.
Zhenghao Lin, Yeyun Gong, Yelong Shen, Zhihao Fan, Chen Lin 0001, Nan Duan 0001, Weizhu Chen
ICML2
2023 Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models
abstract
Large language models can perform various reasoning tasks by using chain-of-thought prompting, which guides them to find answers through step-by-step demonstrations. However, the quality of the prompts depends on the demonstrations given to the models, and creating many of them by hand is costly. We introduce Synthetic prompting, a method that leverages a few handcrafted examples to prompt the model to generate more examples by itself, and selects effective demonstrations to elicit better reasoning. Our method alternates between a backward and forward process to generate new examples. The backward process generates a question that match a sampled reasoning chain, so that the question is solvable and clear. The forward process produces a more detailed reasoning chain for the question, improving the quality of the example. We evaluate our method on numerical, symbolic, and algorithmic reasoning tasks, and show that it outperforms existing prompting techniques.
Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan 0001, Weizhu Chen
ICML2
2023 On-the-Fly Adapting Code Summarization on Trainable Cost-Effective Language Models
abstract
Deep learning models are emerging to summarize source code to comment, facilitating tasks of code documentation and program comprehension. Scaled-up large language models trained on large open corpus have achieved good performance in such tasks. However, in practice, the subject code in one certain project can be specific, which may not align with the overall training corpus. Some code samples from other projects may be contradictory and introduce inconsistencies when the models try to fit all the samples. In this work, we introduce a novel approach, Adacom, to improve the performance of comment generators by on-the-fly model adaptation. This research is motivated by the observation that deep comment generators often need to strike a balance as they need to fit all the training samples. Specifically, for one certain target code $c$, some training samples $S_p$ could have made more contributions while other samples $S_o$ could have counter effects. However, the traditional fine-tuned models need to fit both $S_p$ and $S_o$ from a global perspective, leading to compromised performance for one certain target code $c$. In this context, we design Adacom to (1) detect whether the model might have a compromised performance on a target code $c$ and (2) retrieve a few helpful training samples $S_p$ that have contradictory samples in the training dataset and, (3) adapt the model on the fly by re-training the $S_p$ to strengthen the helpful samples and unlearn the harmful samples. Our extensive experiments on 7 comment generators and 4 public datasets show that (1) can significantly boost the performance of comment generation (BLEU4 score by on average 14.9\%, METEOR by 12.2\%, and ROUGE-L by 7.4\%), (2) the adaptation on one code sample is cost-effective and acceptable as an on-the-fly solution, and (3) can adapt well on out-of-distribution code samples.
Yufan Cai 0001, Yun Lin 0001, Chenyan Liu, Jinglian Wu, Yifan Zhang 0019, Yeyun Gong, Jin Song Dong 0001
NeurIPS7
2023 AR-Diffusion: Auto-Regressive Diffusion Model for Text Generation
abstract
Diffusion models have gained significant attention in the realm of image generation due to their exceptional performance. Their success has been recently expanded to text generation via generating all tokens within a sequence concurrently. However, natural language exhibits a far more pronounced sequential dependency in comparison to images, and the majority of existing language models are trained with a left-to-right auto-regressive approach. To account for the inherent sequential characteristic of natural language, we introduce Auto-Regressive Diffusion (AR-Diffusion). AR-Diffusion ensures that the generation of tokens on the right depends on the generated ones on the left, a mechanism achieved through employing a dynamic number of denoising steps that vary based on token position. This results in tokens on the left undergoing fewer denoising steps than those on the right, thereby enabling them to generate earlier and subsequently influence the generation of tokens on the right. In a series of experiments on various text generation tasks, including text summarization, machine translation, and common sense generation, AR-Diffusion clearly demonstrated its superiority over existing diffusion language models and that it can be $100\times\sim600\times$ faster when achieving comparable results. Our code is available at https://github.com/microsoft/ProphetNet/tree/master/AR-diffusion.
Zhihao Fan, Xiao Liu 0029, Hai-Tao Zheng 0002, Yeyun Gong, Yelong Shen, Jian Jiao 0007, Zhongyu Wei, Jian Guo 0016, Nan Duan 0001, Weizhu Chen
NeurIPS5
2023 MASTER: Multi-task Pre-trained Bottlenecked Masked Autoencoders Are Better Dense Retrievers
Kun Zhou 0002, Xiao Liu 0029, Yeyun Gong, Wayne Xin Zhao, Daxin Jiang, Nan Duan 0001, Ji-Rong Wen
ECML/PKDD (2)3
2023 PROD: Progressive Distillation for Dense Retrieval
abstract
Knowledge distillation is an effective way to transfer knowledge from a strong teacher to an efficient student model. Ideally, we expect the better the teacher is, the better the student performs. However, this expectation does not always come true. It is common that a strong teacher model results in a bad student via distillation due to the nonnegligible gap between teacher and student. To bridge the gap, we propose PROD, a PROgressive Distillation method, for dense retrieval. PROD consists of a teacher progressive distillation and a data progressive distillation to gradually improve the student. To alleviate catastrophic forgetting, we introduce a regularization term in each distillation process. We conduct extensive experiments on seven datasets including five widely-used publicly available benchmarks: MS MARCO Passage, TREC Passage 19, TREC Document 19, MS MARCO Document, and Natural Questions, as well as two industry datasets: Bing-Rel and Bing-Ads. PROD achieves the state-of-the-art in the distillation methods for dense retrieval. Our 6-layer student model even surpasses most of the existing 12-layer models on all five public benchmarks. The code and models are released in https://github.com/microsoft/SimXNS.
Zhenghao Lin, Yeyun Gong, Xiao Liu 0029, Hang Zhang 0029, Chen Lin 0001, Anlei Dong, Jian Jiao 0007, Jingwen Lu, Daxin Jiang, Rangan Majumder, Nan Duan 0001
WWW2
2022 DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response Generation
abstract
Wei Chen, Yeyun Gong, Song Wang, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Yi Mao, Weizhu Chen, Biao Cheng, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Wei Chen 0088, Yeyun Gong, Song Wang 0012, Bolun Yao, Weizhen Qi, Zhongyu Wei, Xiaowu Hu, Bartuer Zhou, Weizhu Chen, Biao Cheng, Nan Duan 0001
ACL (1)2
2022 Contextual Fine-to-Coarse Distillation for Coarse-grained Response Selection in Open-Domain Conversations
abstract
Wei Chen, Yeyun Gong, Can Xu, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Wei Chen 0088, Yeyun Gong, Can Xu 0002, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan 0001
ACL (1)2
2022 Metric-guided Distillation: Distilling Knowledge from the Metric to Ranker and Retriever for Generative Commonsense Reasoning
abstract
Xingwei He, Yeyun Gong, A-Long Jin, Weizhen Qi, Hang Zhang, Jian Jiao, Bartuer Zhou, Biao Cheng, Sm Yiu, Nan Duan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xingwei He 0003, Yeyun Gong, A-Long Jin, Weizhen Qi, Hang Zhang 0029, Jian Jiao 0007, Bartuer Zhou, Biao Cheng, Siu-Ming Yiu, Nan Duan 0001
EMNLP2
2022 Sentiment-Aware Word and Sentence Level Pre-training for Sentiment Analysis
abstract
Most existing pre-trained language representation models (PLMs) are sub-optimal in sentiment analysis tasks, as they capture the sentiment information from word-level while underconsidering sentence-level information.In this paper, we propose SentiWSP, a novel Sentiment-aware pre-trained language model with combined Word-level and Sentence-level Pre-training tasks.The word level pre-training task detects replaced sentiment words, via a generator-discriminator framework, to enhance the PLM's knowledge about sentiment words.The sentence level pre-training task further strengthens the discriminator via a contrastive learning framework, with similar sentences as negative samples, to encode sentiments in a sentence.Extensive experimental results show that SentiWSP achieves new state-of-the-art performance on various sentence-level and aspectlevel sentiment classification benchmarks.We have made our code and model publicly available at https://github.com/XMUDM/SentiWSP.
Shuai Fan 0007, Chen Lin 0001, Haonan Li 0002, Zhenghao Lin, Jinsong Su, Hang Zhang 0029, Yeyun Gong, Jian Guo 0016, Nan Duan 0001
EMNLP7
2022 CodeRetriever: A Large Scale Contrastive Pre-Training Method for Code Search
abstract
Xiaonan Li, Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, Nan Duan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Yeyun Gong, Yelong Shen, Xipeng Qiu, Hang Zhang 0029, Bolun Yao, Weizhen Qi, Daxin Jiang, Weizhu Chen, Nan Duan 0001
EMNLP2
2022 Adversarial Retriever-Ranker for Dense Text Retrieval
Hang Zhang 0029, Yeyun Gong, Yelong Shen, Jiancheng Lv 0001, Nan Duan 0001, Weizhu Chen
ICLR2
2022 Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense Embeddings
abstract
Vector quantization (VQ) based ANN indexes, such as Inverted File System (IVF) and Product Quantization (PQ), have been widely applied to embedding based document retrieval thanks to the competitive time and memory efficiency. Originally, VQ is learned to minimize the reconstruction loss, i.e., the distortions between the original dense embeddings and the reconstructed embeddings after quantization. Unfortunately, such an objective is inconsistent with the goal of selecting ground-truth documents for the input query, which may cause severe loss of retrieval quality. Recent works identify such a defect, and propose to minimize the retrieval loss through contrastive learning. However, these methods intensively rely on queries with ground-truth documents, whose performance is limited by the insufficiency of labeled data. In this paper, we propose Distill-VQ, which unifies the learning of IVF and PQ within a knowledge distillation framework. In Distill-VQ, the dense embeddings are leveraged as "teachers'', which predict the query's relevance to the sampled documents. The VQ modules are treated as the "students'', which are learned to reproduce the predicted relevance, such that the reconstructed embeddings may fully preserve the retrieval result of the dense embeddings. By doing so, Distill-VQ is able to derive substantial training signals from the massive unlabeled data, which significantly contributes to the retrieval quality. We perform comprehensive explorations for the optimal conduct of knowledge distillation, which may provide useful insights for the learning of VQ based ANN index. We also experimentally show that the labeled data is no longer a necessity for high-quality vector quantization, which indicates Distill-VQ's strong applicability in practice. The evaluations are performed on MS MARCO and Natural Questions benchmarks, where Distill-VQ notably outperforms the SOTA VQ methods in Recall and MRR. Our code is avaliable at https://github.com/staoxiao/LibVQ.
Shitao Xiao, Zheng Liu 0011, Weihao Han, Jianjin Zhang, Defu Lian, Yeyun Gong, Qi Chen 0009, Fan Yang 0024, Hao Sun 0015, Yingxia Shao, Xing Xie 0001
SIGIR6
2021 BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale Pretraining
abstract
In this paper, we propose BANG, a new pretraining model to Bridge the gap between Autoregressive (AR) and Non-autoregressive (NAR) Generation. AR and NAR generation can be uniformly regarded as to what extent previous tokens can be attended, and BANG bridges AR and NAR generation through designing a novel model structure for large-scale pre-training. A pretrained BANG model can simultaneously support AR, NAR, and semi-NAR generation to meet different requirements. Experiments on question generation (SQuAD 1.1), summarization (XSum), and dialogue generation (PersonaChat) show that BANG improves NAR and semi-NAR performance significantly as well as attaining comparable performance with strong AR pretrained models. Compared with the semi-NAR strong baselines, BANG achieves absolute improvements of 14.01 and 5.24 in the overall scores of SQuAD 1.1 and XSum, respectively. In addition, BANG achieves absolute improvements of 10.73, 6.39, and 5.90 in the overall scores of SQuAD, XSUM, and PersonaChat compared with the NAR strong baselines, respectively. Our code will be made publicly available.
Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Weizhu Chen, Dayiheng Liu, Kewen Tang, Houqiang Li, Jiusheng Chen, Ruofei Zhang, Ming Zhou 0001, Nan Duan 0001
ICML2
2021 EL-Attention: Memory Efficient Lossless Attention for Generation
abstract
Transformer model with multi-head attention requires caching intermediate results for efficient inference in generation tasks. However, cache brings new memory-related costs and prevents leveraging larger batch size for faster speed. We propose memory-efficient lossless attention (called EL-attention) to address this issue. It avoids heavy operations for building multi-head keys and values, cache for them is not needed. EL-attention constructs an ensemble of attention results by expanding query while keeping key and value shared. It produces the same result as multi-head attention with less GPU memory and faster inference speed. We conduct extensive experiments on Transformer, BART, and GPT-2 for summarization and question generation tasks. The results show EL-attention speeds up existing models by 1.6x to 5.3x without accuracy loss.
Jiusheng Chen, Weizhen Qi, Nikhil Bhendawade, Yeyun Gong, Nan Duan 0001, Ruofei Zhang
ICML5
2021 Poolingformer: Long Document Modeling with Pooling Attention
abstract
In this paper, we introduce a two-level attention schema, Poolingformer, for long document modeling. Its first level uses a smaller sliding window pattern to aggregate information from neighbors. Its second level employs a larger window to increase receptive fields with pooling attention to reduce both computational cost and memory consumption. We first evaluate Poolingformer on two long sequence QA tasks: the monolingual NQ and the multilingual TyDi QA. Experimental results show that Poolingformer sits atop three official leaderboards measured by F1, outperforming previous state-of-the-art models by 1.9 points (79.8 vs. 77.9) on NQ long answer, 1.9 points (79.5 vs. 77.6) on TyDi QA passage answer, and 1.6 points (67.6 vs. 66.0) on TyDi QA minimal answer. We further evaluate Poolingformer on a long sequence summarization task. Experimental results on the arXiv benchmark continue to demonstrate its superior performance.
Hang Zhang 0029, Yeyun Gong, Yelong Shen, Weisheng Li 0001, Jiancheng Lv 0001, Nan Duan 0001, Weizhu Chen
ICML2
2021 Mask Attention Networks: Rethinking and Strengthen Transformer
abstract
Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang, Jian Jiao, Nan Duan, Ruofei Zhang, Xuanjing Huang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zhihao Fan, Yeyun Gong, Dayiheng Liu, Zhongyu Wei, Siyuan Wang 0025, Jian Jiao 0007, Nan Duan 0001, Ruofei Zhang, Xuanjing Huang 0001
NAACL-HLT2
2021 Question Generation from Code Snippets and Programming Error Messages
Bolun Yao, Wei Chen 0088, Yeyun Gong, Bartuer Zhou, Zhongyu Wei, Biao Cheng, Nan Duan 0001
NLPCC (1)3
2020 Graph-Based Transformer with Cross-Candidate Verification for Semantic Parsing
abstract
In this paper, we present a graph-based Transformer for semantic parsing. We separate the semantic parsing task into two steps: 1) Use a sequence-to-sequence model to generate the logical form candidates. 2) Design a graph-based Transformer to rerank the candidates. To handle the structure of logical forms, we incorporate graph information to Transformer, and design a cross-candidate verification mechanism to consider all the candidates in the ranking process. Furthermore, we integrate BERT into our model and jointly train the graph-based Transformer and BERT. We conduct experiments on 3 semantic parsing benchmarks, ATIS, JOBS and Task Oriented semantic Parsing dataset (TOP). Experiments show that our graph-based reranking model achieves results comparable to state-of-the-art models on the ATIS and JOBS datasets. And on the TOP dataset, our model achieves a new state-of-the-art result.
Yeyun Gong, Weizhen Qi, Guihong Cao, Jianshu Ji, Xiaola Lin
AAAI2
2020 Neural Semantic Parsing in Low-Resource Settings with Back-Translation and Meta-Learning
abstract
Neural semantic parsing has achieved impressive results in recent years, yet its success relies on the availability of large amounts of supervised data. Our goal is to learn a neural semantic parser when only prior knowledge about a limited number of simple rules is available, without access to either annotated programs or execution results. Our approach is initialized by rules, and improved in a back-translation paradigm using generated question-program pairs from the semantic parser and the question generator. A phrase table with frequent mapping patterns is automatically derived, also updated as training progresses, to measure the quality of generated instances. We train the model with model-agnostic meta-learning to guarantee the accuracy and stability on examples covered by rules, and meanwhile acquire the versatility to generalize well on examples uncovered by rules. Results on three benchmark datasets with different domains and programs show that our approach incrementally improves the accuracy. On WikiSQL, our best model is comparable to the state-of-the-art system learned from denotations.
Duyu Tang, Nan Duan 0001, Yeyun Gong, Bing Qin 0001, Daxin Jiang
AAAI4
2020 RikiNet: Reading Wikipedia Pages for Natural Question Answering
abstract
Reading long documents to answer opendomain questions remains challenging in natural language understanding.In this paper, we introduce a new model, called RikiNet, which reads Wikipedia pages for natural question answering.RikiNet contains a dynamic paragraph dual-attention reader and a multi-level cascaded answer predictor.The reader dynamically represents the document and question by utilizing a set of complementary attention mechanisms.The representations are then fed into the predictor to obtain the span of the short answer, the paragraph of the long answer, and the answer type in a cascaded manner.On the Natural Questions (NQ) dataset, a single RikiNet achieves 74.3 F1 and 57.9 F1 on longanswer and short-answer tasks.To our best knowledge, it is the first single model that outperforms the single human performance.Furthermore, an ensemble RikiNet obtains 76.1 F1 and 61.3 F1 on long-answer and shortanswer tasks, achieving the best performance on the official NQ leaderboard 1 .
Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Daxin Jiang, Jiancheng Lv 0001, Nan Duan 0001
ACL2
2020 An Enhanced Knowledge Injection Model for Commonsense Generation
abstract
Commonsense generation aims at generating plausible everyday scenario description based on a set of provided concepts. Digging the relationship of concepts from scratch is non-trivial, therefore, we retrieve prototypes from external knowledge to assist the understanding of the scenario for better description generation. We integrate two additional modules into the pretrained encoder-decoder model for prototype modeling to enhance the knowledge injection procedure. We conduct experiment on CommonGen benchmark, experimental results show that our method significantly improves the performance on all the metrics.
Zhihao Fan, Yeyun Gong, Zhongyu Wei, Siyuan Wang 0025, Yameng Huang, Jian Jiao 0007, Xuanjing Huang 0001, Nan Duan 0001, Ruofei Zhang
COLING2
2020 Multi-level Alignment Pretraining for Multi-lingual Semantic Parsing
abstract
In this paper, we present a multi-level alignment pretraining method in a unified architecture for multi-lingual semantic parsing.In this architecture, we use an adversarial training method to align the space of different languages and use sentence level and word level parallel corpus as supervision information to align the semantic of different languages.Finally, we jointly train the multi-level alignment and semantic parsing tasks.We conduct experiments on a publicly available multi-lingual semantic parsing dataset ATIS and a newly constructed dataset.Experimental results show that our model outperforms state-of-the-art methods on both datasets.
Yeyun Gong, Weizhen Qi, Nan Duan 0001, Xiaola Lin
COLING2
2020 Uncertainty-Aware Label Refinement for Sequence Labeling
abstract
Conditional random fields (CRF) for label decoding has become ubiquitous in sequence labeling tasks.However, the local label dependencies and inefficient Viterbi decoding have always been a problem to be solved.In this work, we introduce a novel two-stage label decoding framework to model long-term label dependencies, while being much more computationally efficient.A base model first predicts draft labels, and then a novel twostream self-attention model makes refinements on these draft predictions based on longrange label dependencies, which can achieve parallel decoding for a faster prediction.In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing error propagation.The experimental results on three sequence labeling benchmarks demonstrated that the proposed method not only outperformed the CRF-based methods but also greatly accelerated the inference process.* Both authors contributed equally.
Tao Gui, Jiacheng Ye, Qi Zhang 0001, Zhengyan Li, Zichu Fei, Yeyun Gong, Xuanjing Huang 0001
EMNLP (1)6
2020 XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and Generation
abstract
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Daniel Campos, Rangan Majumder, Ming Zhou. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Yaobo Liang, Nan Duan 0001, Yeyun Gong, Ning Wu 0013, Fenfei Guo, Weizhen Qi, Ming Gong 0001, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Dong Bo Cui, Sining Wei, Taroon Bharti, Jiun-Hung Chen, Winnie Wu, Fan Yang 0024, Daniel Campos, Rangan Majumder, Ming Zhou 0001
EMNLP (1)3
2020 Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous Space
abstract
In this paper, we propose a novel data augmentation method, referred to as Controllable Rewriting based Question Data Augmentation (CRQDA), for machine reading comprehension (MRC), question generation, and question-answering natural language inference tasks.We treat the question data augmentation task as a constrained question rewriting problem to generate context-relevant, high-quality, and diverse question data samples.CRQDA utilizes a Transformer autoencoder to map the original discrete question into a continuous embedding space.It then uses a pre-trained MRC model to revise the question representation iteratively with gradientbased optimization.Finally, the revised question representations are mapped back into the discrete space, which serve as additional question data.Comprehensive experiments on SQuAD 2.0, SQuAD 1.1 question generation, and QNLI tasks demonstrate the effectiveness of CRQDA 1 .
Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Jiusheng Chen, Jiancheng Lv 0001, Nan Duan 0001, Ming Zhou 0001
EMNLP (1)2
2020 Diverse, Controllable, and Keyphrase-Aware: A Corpus and Method for News Multi-Headline Generation
abstract
News headline generation aims to produce a short sentence to attract readers to read the news.One news article often contains multiple keyphrases that are of interest to different users, which can naturally have multiple reasonable headlines.However, most existing methods focus on the single headline generation.In this paper, we propose generating multiple headlines with keyphrases of user interests, whose main idea is to generate multiple keyphrases of interest to users for the news first, and then generate multiple keyphrase-relevant headlines.We propose a multi-source Transformer decoder, which takes three sources as inputs: (a) keyphrase, (b) keyphrase-filtered article, and (c) original article to generate keyphrase-relevant, highquality, and diverse headlines.Furthermore, we propose a simple and effective method to mine the keyphrases of interest in the news article and build a first large-scale keyphraseaware news headline corpus, which contains over 180K aligned triples of news article, headline, keyphrase .Extensive experimental comparisons on the real-world dataset show that the proposed method achieves state-of-theart results in terms of quality and diversity 1 .
Dayiheng Liu, Yeyun Gong, Jie Fu 0001, Daxin Jiang, Jiancheng Lv 0001, Nan Duan 0001
EMNLP (1)2
2020 Leveraging Document-Level Label Consistency for Named Entity Recognition
abstract
Document-level label consistency is an effective indicator that different occurrences of a particular token sequence are very likely to have the same entity types. Previous work focused on better context representations and used the CRF for label decoding. However, CRF-based methods are inadequate for modeling document-level label consistency. This work introduces a novel two-stage label refinement approach to handle document-level label consistency, where a key-value memory network is first used to record draft labels predicted by the base model, and then a multi-channel Transformer makes refinements on these draft predictions based on the explicit co-occurrence relationship derived from the memory network. In addition, in order to mitigate the side effects of incorrect draft labels, Bayesian neural networks are used to indicate the labels with a high probability of being wrong, which can greatly assist in preventing the incorrect refinement of correct draft labels. The experimental results on three named entity recognition benchmarks demonstrated that the proposed method significantly outperformed the state-of-the-art methods.
Tao Gui, Jiacheng Ye, Qi Zhang 0001, Yaqian Zhou 0001, Yeyun Gong, Xuanjing Huang 0001
IJCAI5
2020 ProphetNet-Ads: A Looking Ahead Strategy for Generative Retrieval Models in Sponsored Search Engine
Weizhen Qi, Yeyun Gong, Jian Jiao 0007, Ruofei Zhang, Houqiang Li, Nan Duan 0001, Ming Zhou 0001
NLPCC (2)2
2019 Joint Type Inference on Entities and Relations via Graph Convolutional Networks
abstract
We develop a new paradigm for the task of joint entity relation extraction.It first identifies entity spans, then performs a joint inference on entity types and relation types.To tackle the joint type inference task, we propose a novel graph convolutional network (GCN) running on an entity-relation bipartite graph.By introducing a binary relation classification task, we are able to utilize the structure of entity-relation bipartite graph in a more efficient and interpretable way.Experiments on ACE05 show that our model outperforms existing joint models in entity performance and is competitive with the state-of-the-art in relation performance.
Changzhi Sun, Yeyun Gong, Yuanbin Wu, Ming Gong 0001, Daxin Jiang, Man Lan, Shiliang Sun, Nan Duan 0001
ACL (1)2
2019 Aggregating Bidirectional Encoder Representations Using MatchLSTM for Sequence Matching
abstract
Bo Shao, Yeyun Gong, Weizhen Qi, Nan Duan, Xiaola Lin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yeyun Gong, Weizhen Qi, Nan Duan 0001, Xiaola Lin
EMNLP/IJCNLP (1)2
2019 Weakly Supervised Multi-task Learning for Semantic Parsing
abstract
Semantic parsing is a challenging and important task which aims to convert a natural language sentence to a logical form. Existing neural semantic parsing methods mainly use (Q-L) pairs to train a sequence-to-sequence model. However, the amount of existing Q-L labeled data is limited and hard to obtain. We propose an effective method which substantially utilizes labeling information from other tasks to enhance the training of a semantic parser. We design a multi-task learning model to train question type classification, entity mention detection together with question semantic parsing using a shared encoder. We propose a weakly supervised learning method to enhance our multi-task learning model with paraphrase data, based on the idea that the paraphrased questions should have the same logical form and question type information. Finally, we integrate the weakly supervised multi-task learning method to an encoder-decoder framework. Experiments on a newly constructed dataset and ComplexWebQuestions show that our proposed method outperforms state-of-the-art methods which demonstrates the effectiveness and robustness of our method.
Yeyun Gong, Junwei Bao 0001, Jianshu Ji, Guihong Cao, Xiaola Lin, Nan Duan 0001
IJCAI2
2018 Hashtag recommendation for multimodal microblog posts
Yeyun Gong, Qi Zhang 0001, Xuanjing Huang 0001
Neurocomputing1
2018 Question Generation With Doubly Adversarial Nets
abstract
We study the problem of question generation on a specific domain, where there are no labeled data. To address this problem, we propose a novel neural question generation approach called DoubAN, or doubly adversarial nets, which fully utilizes labeled data from other domains (source domains) and unlabeled data from the target domain. Learning a DoubAN involves two adversarial procedures between a question generator and two adversaries. One adversary is a domain-classification discriminator (DC-Dis), which is designed to help the generator learn domain-general representations of the input text. The other is a question-answering discriminator (QA-Dis), which provides more training data with estimated reward scores for generated text-question pairs. We conduct experiments on the SQuAD dataset as target-domain unlabeled data and the NewsQA dataset as source-domain labeled data. Experiment results show that our DoubAN achieves better results than baselines. Compared to model variants, which adopt only DC-Dis or QA-Dis, we find that the DC-Dis and QA-Dis indirectly interact with each other and jointly improve the quality of generated questions on the target domain. Moreover, extensive analysis and discussion prove the reasonableness and effectiveness of our proposed approach.
Junwei Bao 0001, Yeyun Gong, Nan Duan 0001, Ming Zhou 0001, Tiejun Zhao
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Hashtag Recommendation for Multimodal Microblog Using Co-Attention Network
abstract
In microblogging services, authors can use hashtags to mark keywords or topics. Many live social media applications (e.g., microblog retrieval, classification) can gain great benefits from these manually labeled tags. However, only a small portion of microblogs contain hashtags inputed by users. Moreover, many microblog posts contain not only textual content but also images. These visual resources also provide valuable information that may not be included in the textual content. So that it can also help to recommend hashtags more accurately. Motivated by the successful use of the attention mechanism, we propose a co-attention network incorporating textual and visual information to recommend hashtags for multimodal tweets. Experimental result on the data collected from Twitter demonstrated that the proposed method can achieve better performance than state-of-the-art methods using textual information only.
Qi Zhang 0001, Haoran Huang, Xuanjing Huang 0001, Yeyun Gong
IJCAI5
2017 Hierarchical Dirichlet Processes with Social Influence
Yeyun Gong, Qi Zhang 0001, Xuanjing Huang 0001
NLPCC2
2017 Phrase-based hashtag recommendation for microblog posts
Yeyun Gong, Qi Zhang 0001, Xiaoying Han, Xuanjing Huang 0001
Sci. China Inf. Sci.1
2016 Retweet Prediction with Attention-based Deep Neural Network
abstract
On Twitter-like social media sites, the re-posting statuses or tweets of other users are usually considered to be the key mechanism for spreading information. How to predict whether a tweet will be retweeted by a user has received increasing attention in recent years. Previous methods studied the problem using various linguistic features, personal information of users, and many other manually constructed features to achieve the task. Usually, feature engineering is a laborious task, we require to obtain the external sources and they are difficult or not always available. Recently, deep learning methods have been used in the industry and research community for their ability to learn optimal features automatically and in many tasks, deep learning methods can achieve state-of-the art performance, such as natural language processing, computer vision, image classification and so on. In this work, we proposed a novel attention-based deep neural network to incorporate contextual and social information for this task. We used embeddings to represent the user, the user's attention interests, the author and tweet respectively. To train and evaluate the proposed methods, we also constructed a large dataset collected from Twitter. Experimental results showed that the proposed method could achieve better results than the previous state-of-the-art methods.
Qi Zhang 0001, Yeyun Gong, Jindou Wu, Haoran Huang, Xuanjing Huang 0001
CIKM2
2016 Hashtag Recommendation Using End-To-End Memory Networks with Hierarchical Attention
abstract
On microblogging services, people usually use hashtags to mark microblogs, which have a specific theme or content, making them easier for users to find. Hence, how to automatically recommend hashtags for microblogs has received much attention in recent years. Previous deep neural network-based hashtag recommendation approaches converted the task into a multi-class classification problem. However, most of these methods only took the microblog itself into consideration. Motivated by the intuition that the history of users should impact the recommendation procedure, in this work, we extend end-to-end memory networks to perform this task. We incorporate the histories of users into the external memory and introduce a hierarchical attention mechanism to select more appropriate histories. To train and evaluate the proposed method, we also construct a dataset based on microblogs collected from Twitter. Experimental results demonstrate that the proposed methods can significantly outperform state-of-the-art methods. By incorporating the hierarchical attention mechanism, the relative improvement in the proposed method over the state-of-the-art method is around 67.9% in the F1-score.
Haoran Huang, Qi Zhang 0001, Yeyun Gong, Xuanjing Huang 0001
COLING3
2016 Keyphrase Extraction Using Deep Recurrent Neural Networks on Twitter
abstract
Keyphrases can provide highly condensed and valuable information that allows users to quickly acquire the main ideas.The task of automatically extracting them have received considerable attention in recent decades.Different from previous studies, which are usually focused on automatically extracting keyphrases from documents or articles, in this study, we considered the problem of automatically extracting keyphrases from tweets.Because of the length limitations of Twitter-like sites, the performances of existing methods usually drop sharply.We proposed a novel deep recurrent neural network (RNN) model to combine keywords and context information to perform this problem.To evaluate the proposed method, we also constructed a large-scale dataset collected from Twitter.The experimental results showed that the proposed method performs significantly better than previous methods.
Qi Zhang 0001, Yang Wang 0091, Yeyun Gong, Xuanjing Huang 0001
EMNLP3
2015 Retweet Behavior Prediction Using Hierarchical Dirichlet Process
abstract
The task of predicting retweet behavior is an important and essential step for various social network applications, such as business intelligence, popular event prediction, and so on. Due to the increasing requirements, in recent years, the task has attracted extensive attentions. In this work, we propose a novel method using non-parametric statistical models to combine structural, textual, and temporal information together to predict retweet behavior. To evaluate the proposed method, we collect a large number of microblogs and their corresponding social networks from a real microblog service. Experimental results on the constructed dataset demonstrate that the proposed method can achieve better performance than state-of-the-art methods. The relative improvement of the the proposed over the method using only textual information is more than 38.5% in terms of F1-Score.
Qi Zhang 0001, Yeyun Gong, Xuanjing Huang 0001
AAAI2
2015 Who Will You "@"?
abstract
In Twitter-like social networking services, people can use the "@" symbol to mention other users in tweets and send them a message or link to their profiles. In recent years, social media services are rapidly growing with thousands of millions of users participating in them every day. When the "@" symbol is entered, there should be an automatic suggestion function which recommends a small list of candidates in order to help users to easily identify and input usernames. In this paper, we present our work on building a recommendation system for the mention function in microblogging services. The recommendation strategy we used takes into consideration not only content of the microblog but also histories of candidate users. To better handle these textual information, we propose a novel method that extends the translation-based model. Experimental results on the dataset we collected from a real world microblogging service demonstrate that the proposed method outperforms state-of-the-art approaches.
Yeyun Gong, Qi Zhang 0001, Xuyang Sun, Xuanjing Huang 0001
CIKM1
2015 Hashtag Recommendation Using Dirichlet Process Mixture Models Incorporating Types of Hashtags
abstract
In recent years, the task of recommending hashtags for microblogs has been given increasing attention.Various methods have been proposed to study the problem from different aspects.However, most of the recent studies have not considered the differences in the types or uses of hashtags.In this paper, we introduce a novel nonparametric Bayesian method for this task.Based on the Dirichlet Process Mixture Models (DPMM), we incorporate the type of hashtag as a hidden variable.The results of experiments on the data collected from a real world microblogging service demonstrate that the proposed method outperforms stateof-the-art methods that do not consider these aspects.By taking these aspects into consideration, the relative improvement of the proposed method over the state-of-theart methods is around 12.2% in F1-score.
Yeyun Gong, Qi Zhang 0001, Xuanjing Huang 0001
EMNLP1
2014 A Generative Model for Identifying Target Companies of Microblogs
Yeyun Gong, Yaqian Zhou 0001, Qi Zhang 0001, Xuanjing Huang 0001
COLING1
2014 Time-aware Personalized Hashtag Recommendation on Social Media
Qi Zhang 0001, Yeyun Gong, Xuyang Sun, Xuanjing Huang 0001
COLING2
2013 Map search via a factor graph model
abstract
Map search has received considerable attention in recent years. With map search, users can specify target locations with textual queries. However, these queries do not always include well-formed addresses or place names. They may contain transpositions, misspellings, fragments and so on. Queries may significantly differ from items stored in the spatial database. In this paper, we propose to connect this task to the semi-structured retrieval problem. A novel factor graph-based semi-structured retrieval framework is introduced to incorporate concept weighting, attribute selection, and word-based similarity metrics together. We randomly sampled a number of queries from logs of a commercial map search engine and manually labeled their categories and relevant results for analysis and evaluation. The results of several experimental comparisons demonstrate that our method outperforms both state-of-the-art semi-structured retrieval methods and some commercial systems in retrieving freeform location queries.
Qi Zhang 0001, Jihua Kang, Yeyun Gong, Huan Chen 0012, Yaqian Zhou 0001, Xuanjing Huang 0001
CIKM3
2013 Detecting Spammers in Community Question Answering
Zhuoye Ding, Yeyun Gong, Yaqian Zhou 0001, Qi Zhang 0001, Xuanjing Huang 0001
IJCNLP2
2011 Classical Mongolian Words Recognition in Historical Document
abstract
There are many classical Mongolian historical documents which are reserved in image form, and as a result it is difficult for us to explore and retrieve them. In this paper, we investigate the peculiarities of classical Mongolian documents and propose an approach to recognize the words in them. We design an algorithm to segment the Mongolian words into several Glyph Units(Glyph Unit abbr. GU). Each GU is consisted of no more than three characters. Then we used a three-stage method to recognize the GUs. At the first stage, all the GUs are classified into nine groups by decision tree using three features of the GUs. At the second stage, the GUs in each group are classified individually by five independent BP Neutral Networks whose inputs are other five feature vectors of the GUs. At the last stage, the five results of each GU group from the above five classifiers are combined to provide the final recognized result. The recognition rate of the Mongolian words in our experiment achieves 71%, indicating that our method is effective.
Guanglai Gao, Xiangdong Su, Hongxi Wei, Yeyun Gong
ICDAR4