Tao Ge 0001

dblp:136/7923 · DBLP profile ↗
← Back
43ranked-venue papers
14as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 42 · 13 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
YearPublicationVenuePosition
2025 Low-Bit Quantization Favors Undertrained LLMs
abstract
Low-bit quantization improves machine learning model efficiency but surprisingly favors undertrained large language models (LLMs).Larger models or those trained on fewer tokens exhibit less quantization-induced degradation (QiD), while smaller, well-trained models face significant performance losses.To gain deeper insights into this trend, we study over 1500+ quantized LLM checkpoints of various sizes and at different training levels (undertrained or fully trained) in a controlled setting, deriving scaling laws for understanding the relationship between QiD and factors: the number of training tokens, model size and bit width.With our derived scaling laws, we propose a novel perspective that we can use QiD to measure an LLM's training levels and determine the number of training tokens required for fully training LLMs of various sizes.Moreover, we use the scaling laws to predict the quantization performance of different-sized LLMs trained with 100 trillion tokens.Our projection shows that the low-bit quantization performance of future models, which are expected to be trained with over 100 trillion tokens, may NOT be desirable.This poses a potential challenge for low-bit quantization in the future and highlights the need for awareness of a model's training level when evaluating lowbit quantization research.To facilitate future research on this problem, we release all the 1500+ quantized checkpoints used in this work at https://huggingface.co/Xu-Ouyang.
Xu Ouyang, Tao Ge 0001, Thomas Hartvigsen, Zhisong Zhang, Haitao Mi, Dong Yu 0001
ACL (1)2
2025 Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory
abstract
Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks. However, they utilize non-parametric memory as static storage, which lacks learning capability and remains disconnected from the internal information flow of the parametric models, limiting scalability and efficiency. Based on recent interpretability theories of LMs, we reconceptualize the non-parametric memory represented by kNN-LM as a learnable Mixture-of-Neighbors Induction Memory (MoNIM), which synergizes the induction capabilities of attention heads with the memorization strength of feed-forward networks (FFN). By integrating into the model’s information flow, MoNIM functions as an FFN-like bypass layer within the Transformer architecture, enabling effective learning of new knowledge. Extensive experiments demonstrate that MoNIM is a retentive and scalable continual learner in both data- and model-wise, enhancing the scalability and continual learning performance of semiparametric LMs.
Guangyue Peng, Tao Ge 0001, Wen Luo 0001, Wei Li 0101, Houfeng Wang
ACL (1)2
2025 ALYMPICS: LLM Agents Meet Game Theory
abstract
Game theory is a branch of mathematics that studies strategic interactions among rational agents. We propose Alympics (Olympics for Agents), a systematic framework utilizing Large Language Model (LLM) agents for empirical game theory research. Alympics creates a versatile platform for studying complex game theory problems, bridging the gap between theoretical game theory and empirical investigations by providing a controlled environment for simulating human-like strategic interactions with LLM agents. In our pilot case study, the “Water Allocation Challenge”, we explore Alympics through a challenging strategic game focused on the multi-round auction of scarce survival resources. This study demonstrates the framework’s ability to qualitatively and quantitatively analyze game determinants, strategies, and outcomes. Additionally, we conduct a comprehensive human assessment and an in-depth evaluation of LLM agents in rational strategic decision-making scenarios. Our findings highlight LLM agents’ potential to advance game theory knowledge and expand the understanding of their proficiency in emulating human strategic behavior.
Shaoguang Mao, Yuzhe Cai, Yan Xia 0005, Wenshan Wu, Xun Wang 0012, Qiang Guan, Tao Ge 0001, Furu Wei
COLING8
2025 Router-Tuning: A Simple and Effective Approach for Dynamic Depth
abstract
The Mixture of Depths (MoD) was introduced to improve computational efficiency by dynamically skipping less important layers, reducing redundant computation while maintaining model capacity.Despite its promise, existing MoD approaches remain under-explored and face two main challenges: (1) high training costs due to the need to train the entire model along with the routers that determine which layers to skip, and (2) performance degradation when important layers are bypassed.In response to the first issue, we propose Router-Tuning, which fine-tunes only the routers on a small dataset, drastically reducing the computational overhead associated with full model training.For the second challenge, we investigate Router-Tuning across different architectures and granularities, demonstrating its effectiveness on Attention layers and MoE layers.This method preserves the model's performance while significantly enhancing computational and memory efficiency.Extensive experiments demonstrate that our approach delivers competitive results while dramatically improving the computation efficiency, e.g., 21% speedup and only a 0.2% performance drop.
Shwai He, Tao Ge 0001, Guoheng Sun, Bowei Tian, Xiaoyang Wang 0001, Dong Yu 0001
EMNLP2
2025 K-Level Reasoning: Establishing Higher Order Beliefs in Large Language Models for Strategic Reasoning
abstract
Yadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang, Yan Xia, Man Lan, Furu Wei. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Shaoguang Mao, Tao Ge 0001, Xun Wang 0012, Yan Xia 0005, Man Lan, Furu Wei
NAACL (Long Papers)3
2025 Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
abstract
Reinforcement learning from human feedback (RLHF) has demonstrated remarkable effectiveness in aligning large language models (LLMs) with human preferences. Many existing alignment approaches rely on the Bradley-Terry (BT) model assumption, which assumes the existence of a ground-truth reward for each prompt-response pair. However, this assumption can be overly restrictive when modeling complex human preferences. In this paper, we drop the BT model assumption and study LLM alignment under general preferences, formulated as a two-player game. Drawing on theoretical insights from learning in games, we integrate optimistic online mirror descent into our alignment framework to approximate the Nash policy. Theoretically, we demonstrate that our approach achieves an $\mathcal{O}(T^{-1})$ bound on the duality gap, improving upon the previous $\mathcal{O}(T^{-1/2})$ result. Meanwhile, it enjoys a linear convergence rate in the last iterate, a property not achieved by previous methods. More importantly, we implement our method and show through experiments that it outperforms state-of-the-art RLHF algorithms across multiple representative benchmarks.
Dian Yu 0001, Tao Ge 0001, Linfeng Song, Zhichen Zeng 0001, Haitao Mi, Nan Jiang 0008, Dong Yu 0001
NeurIPS3
2024 Learn Beyond The Answer: Training Language Models with Reflection for Mathematical Reasoning
abstract
Supervised fine-tuning enhances the problemsolving abilities of language models across various mathematical reasoning tasks.To maximize such benefits, existing research focuses on broadening the training set with various data augmentation techniques, which is effective for standard single-round question-answering settings.Our work introduces a novel technique aimed at cultivating a deeper understanding of the training problems at hand, enhancing performance not only in standard settings but also in more complex scenarios that require reflective thinking.Specifically, we propose reflective augmentation, a method that embeds problem reflection into each training instance.It trains the model to consider alternative perspectives and engage with abstractions and analogies, thereby fostering a thorough comprehension through reflective reasoning.Extensive experiments validate the achievement of our aim, underscoring the unique advantages of our method and its complementary nature relative to existing augmentation techniques. 1 Question Answer Question
Zhihan Zhang 0001, Tao Ge 0001, Zhenwen Liang, Wenhao Yu 0002, Dian Yu 0001, Mengzhao Jia, Dong Yu 0001, Meng Jiang 0001
EMNLP2
2024 In-context Autoencoder for Context Compression in a Large Language Model
abstract
We propose the In-context Autoencoder (ICAE), leveraging the power of a large language model (LLM) to compress a long context into short compact memory slots that can be directly conditioned on by the LLM for various purposes. ICAE is first pretrained using both autoencoding and language modeling objectives on massive text data, enabling it to generate memory slots that accurately and comprehensively represent the original context. Then, it is fine-tuned on instruction data for producing desirable responses to various prompts. Experiments demonstrate that our lightweight ICAE, introducing about 1% additional parameters, effectively achieves $4\times$ context compression based on Llama, offering advantages in both improved latency and GPU memory cost during inference, and showing an interesting insight in memorization as well as potential for scalability. These promising results imply a novel perspective on the connection between working memory in cognitive science and representation learning in LLMs, revealing ICAE's significant implications in addressing the long context problem and suggesting further research in LLM context management. Our data, code and models are available at https://github.com/getao/icae.
Tao Ge 0001, Jing Hu 0001, Xun Wang 0012, Furu Wei
ICLR1
2024 Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration
abstract
Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge, Furu Wei, Heng Ji. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhenhailong Wang, Shaoguang Mao, Wenshan Wu, Tao Ge 0001, Furu Wei, Heng Ji 0001
NAACL-HLT4
2024 xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
abstract
This paper introduces xRAG, an innovative context compression method tailored for retrieval-augmented generation. xRAG reinterprets document embeddings in dense retrieval--traditionally used solely for retrieval--as features from the retrieval modality. By employing a modality fusion methodology, xRAG seamlessly integrates these embeddings into the language model representation space, effectively eliminating the need for their textual counterparts and achieving an extreme compression rate. In xRAG, the only trainable component is the modality bridge, while both the retriever and the language model remain frozen. This design choice allows for the reuse of offline-constructed document embeddings and preserves the plug-and-play nature of retrieval augmentation. Experimental results demonstrate that xRAG achieves an average improvement of over 10% across six knowledge-intensive tasks, adaptable to various language model backbones, ranging from a dense 7B model to an 8x7B Mixture of Experts configuration. xRAG not only significantly outperforms previous context compression methods but also matches the performance of uncompressed models on several datasets, while reducing overall FLOPs by a factor of 3.53. Our work pioneers new directions in retrieval-augmented generation from the perspective of multimodality fusion, and we hope it lays the foundation for future efficient and scalable retrieval-augmented systems.
Xin Cheng 0002, Xun Wang 0012, Xingxing Zhang 0002, Tao Ge 0001, Furu Wei, Huishuai Zhang, Dongyan Zhao 0001
NeurIPS4
2024 Overview of the NLPCC 2024 Shared Task: Chinese Essay Discourse Logic Evaluation and Integration
Hongyi Wu, Xinshu Shen, Man Lan, Yuanbin Wu, Xiaopeng Bai, Shaoguang Mao, Tao Ge 0001, Yan Xia 0005
NLPCC (5)8
2023 Extensible Prompts for Language Models on Zero-shot Language Style Customization
abstract
We propose eXtensible Prompt (X-Prompt) for prompting a large language model (LLM) beyond natural language (NL). X-Prompt instructs an LLM with not only NL but also an extensible vocabulary of imaginary words. Registering new imaginary words allows us to instruct the LLM to comprehend concepts that are difficult to describe with NL words, thereby making a prompt more descriptive. Also, these imaginary words are designed to be out-of-distribution (OOD) robust so that they can be (re)used like NL words in various prompts, distinguishing X-Prompt from soft prompt that is for fitting in-distribution data. We propose context-augmented learning (CAL) to learn imaginary words for general usability, enabling them to work properly in OOD (unseen) prompts. We experiment X-Prompt for zero-shot language style customization as a case study. The promising results of X-Prompt demonstrate its potential to facilitate advanced interaction beyond the natural language interface, bridging the communication gap between humans and LLMs.
Tao Ge 0001, Jing Hu 0001, Li Dong 0004, Shaoguang Mao, Yan Xia 0005, Xun Wang 0012, Furu Wei
NeurIPS1
2023 Enhancing Detailed Feedback to Chinese Writing Learners Using a Soft-Label Driven Approach and Tag-Aware Ranking Model
Yuzhe Cai, Shaoguang Mao, Chenshuo Wang, Tao Ge 0001, Wenshan Wu, Yan Xia 0005, Chanjin Zheng, Qiang Guan
NLPCC (1)4
2023 Overview of the NLPCC 2023 Shared Task: Chinese Essay Discourse Coherence Evaluation
Hongyi Wu, Xinshu Shen, Man Lan, Xiaopeng Bai, Yuanbin Wu, Aimin Zhou, Shaoguang Mao, Tao Ge 0001, Yan Xia 0005
NLPCC (3)8
2022 Text Revision By On-the-Fly Representation Optimization
abstract
Text revision refers to a family of natural language generation tasks, where the source and target sequences share moderate resemblance in surface form but differentiate in attributes, such as text formality and simplicity. Current state-of-the-art methods formulate these tasks as sequence-to-sequence learning problems, which rely on large-scale parallel training corpus. In this paper, we present an iterative in-place editing approach for text revision, which requires no parallel data. In this approach, we simply fine-tune a pre-trained Transformer with masked language modeling and attribute classification. During inference, the editing at each iteration is realized by two-step span replacement. At the first step, the distributed representation of the text optimizes on the fly towards an attribute function. At the second step, a text span is masked and another new one is proposed conditioned on the optimized representation. The empirical experiments on two typical and important text revision tasks, text formalization and text simplification, show the effectiveness of our approach. It achieves competitive and even better performance than state-of-the-art supervised methods on text simplification, and gains better performance than strong unsupervised methods on text formalization. Our code and model are released at https://github.com/jingjingli01/OREO.
Jingjing Li 0007, Zichao Li 0003, Tao Ge 0001, Irwin King, Michael R. Lyu
AAAI3
2022 EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq Generation
abstract
We introduce EDGEFORMER -a parameterefficient Transformer for on-device seq2seq generation under the strict computation and memory constraints.Compared with the previous parameter-efficient Transformers, EDGE-FORMER applies two novel principles for costeffective parameterization, allowing it to perform better given the same parameter budget; moreover, EDGEFORMER is further enhanced by layer adaptation innovation that is proposed for improving the network with shared layers.Extensive experiments show EDGEFORMER can effectively outperform previous parameterefficient Transformer baselines and achieve competitive results under both the computation and memory constraints.Given the promising results, we release EDGELM 1 -the pretrained version of EDGEFORMER, which is the first publicly available pretrained on-device seq2seq model that can be easily fine-tuned for seq2seq tasks with strong results, facilitating on-device seq2seq generation in practice.
Tao Ge 0001, Furu Wei
EMNLP1
2022 A Unified Strategy for Multilingual Grammatical Error Correction with Pre-trained Cross-Lingual Language Model
abstract
Synthetic data construction of Grammatical Error Correction (GEC) for non-English languages relies heavily on human-designed and language-specific rules, which produce limited error-corrected patterns. In this paper, we propose a generic and language-independent strategy for multilingual GEC, which can train a GEC system effectively for a new non-English language with only two easy-to-access resources: 1) a pre-trained cross-lingual language model (PXLM) and 2) parallel translation data between English and the language. Our approach creates diverse parallel GEC data without any language-specific operations by taking the non-autoregressive translation generated by PXLM and the gold translation as error-corrected sentence pairs. Then, we reuse PXLM to initialize the GEC model and pre-train it with the synthetic data generated by itself, which yields further improvement. We evaluate our approach on three public benchmarks of GEC in different languages. It achieves the state-of-the-art results on the NLPCC 2018 Task 2 dataset (Chinese) and obtains competitive performance on Falko-Merlin (German) and RULEC-GEC (Russian). Further analysis demonstrates that our data construction method is complementary to rule-based approaches.
Xin Sun 0013, Tao Ge 0001, Shuming Ma, Jingjing Li 0007, Furu Wei, Houfeng Wang
IJCAI2
2021 Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding
abstract
Xin Sun, Tao Ge, Furu Wei, Houfeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xin Sun 0013, Tao Ge 0001, Furu Wei, Houfeng Wang
ACL/IJCNLP (1)2
2021 Beyond Preserved Accuracy: Evaluating Loyalty and Robustness of BERT Compression
abstract
Recent studies on compression of pretrained language models (e.g., BERT) usually use preserved accuracy as the metric for evaluation.In this paper, we propose two new metrics, label loyalty and probability loyalty that measure how closely a compressed model (i.e., student) mimics the original model (i.e., teacher).We also explore the effect of compression with regard to robustness under adversarial attacks.We benchmark quantization, pruning, knowledge distillation and progressive module replacing with loyalty and robustness.By combining multiple compression techniques, we provide a practical strategy to achieve better accuracy, loyalty and robustness. 1
Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Julian J. McAuley, Furu Wei
EMNLP (1)3
2021 Improving Sequence-to-Sequence Pre-training via Sequence Span Rewriting
abstract
In this paper, we propose Sequence Span Rewriting (SSR), a self-supervised task for sequence-to-sequence (Seq2Seq) pre-training.SSR learns to refine the machine-generated imperfect text spans into ground truth text.SSR provides more fine-grained and informative supervision in addition to the original textinfilling objective.Compared to the prevalent text infilling objectives for Seq2Seq pretraining, SSR is naturally more consistent with many downstream generation tasks that require sentence rewriting (e.g., text summarization, question generation, grammatical error correction, and paraphrase generation).We conduct extensive experiments by using SSR to improve the typical Seq2Seq pre-trained model T5 in a continual pre-training setting and show substantial improvements over T5 on various natural language generation tasks. 1
Wangchunshu Zhou, Tao Ge 0001, Canwen Xu, Ke Xu 0001, Furu Wei
EMNLP (1)2
2021 Blow the Dog Whistle: A Chinese Dataset for Cant Understanding with Common Sense and World Knowledge
abstract
Canwen Xu, Wangchunshu Zhou, Tao Ge, Ke Xu, Julian McAuley, Furu Wei. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Julian J. McAuley, Furu Wei
NAACL-HLT3
2020 Fact-Aware Sentence Split and Rephrase with Permutation Invariant Training
abstract
Sentence Split and Rephrase aims to break down a complex sentence into several simple sentences with its meaning preserved. Previous studies tend to address the issue by seq2seq learning from parallel sentence pairs, which takes a complex sentence as input and sequentially generates a series of simple sentences. However, the conventional seq2seq learning has two limitations for this task: (1) it does not take into account the facts stated in the long sentence; As a result, the generated simple sentences may miss or inaccurately state the facts in the original sentence. (2) The order variance of the simple sentences to be generated may confuse the seq2seq model during training because the simple sentences derived from the long source sentence could be in any order.To overcome the challenges, we first propose the Fact-aware Sentence Encoding, which enables the model to learn facts from the long sentence and thus improves the precision of sentence split; then we introduce Permutation Invariant Training to alleviate the effects of order variance in seq2seq learning for this task. Experiments on the WebSplit-v1.0 benchmark dataset show that our approaches can largely improve the performance over the previous seq2seq learning approaches. Moreover, an extrinsic evaluation on oie-benchmark verifies the effectiveness of our approaches by an observation that splitting long sentences with our state-of-the-art model as preprocessing is helpful for improving OpenIE performance.
Yinuo Guo, Tao Ge 0001, Furu Wei
AAAI2
2020 Parallel Data Augmentation for Formality Style Transfer
abstract
The main barrier to progress in the task of Formality Style Transfer is the inadequacy of training data.In this paper, we study how to augment parallel data and propose novel and simple data augmentation methods for this task to obtain useful sentence pairs with easily accessible models and systems.Experiments demonstrate that our augmented parallel data largely helps improve formality style transfer when it is used to pre-train the model, leading to the state-of-the-art results in the GYAFC benchmark dataset 1 .
Yi Zhang 0050, Tao Ge 0001, Xu Sun 0001
ACL2
2020 Improving the Efficiency of Grammatical Error Correction with Erroneous Span Detection and Correction
abstract
We propose a novel language-independent approach to improve the efficiency for Grammatical Error Correction (GEC) by dividing the task into two subtasks: Erroneous Span Detection (ESD) and Erroneous Span Correction (ESC).ESD identifies grammatically incorrect text spans with an efficient sequence tagging model.Then, ESC leverages a seq2seq model to take the sentence with annotated erroneous spans as input and only outputs the corrected text for these spans.Experiments show our approach performs comparably to conventional seq2seq approaches in both English and Chinese GEC benchmarks with less than 50% time cost for inference.
Mengyun Chen, Tao Ge 0001, Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001
EMNLP (1)2
2020 BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
abstract
In this paper, we propose a novel model compression approach to effectively compress BERT by progressive module replacing.Our approach first divides the original BERT into several modules and builds their compact substitutes.Then, we randomly replace the original modules with their substitutes to train the compact modules to mimic the behavior of the original modules.We progressively increase the probability of replacement through the training.In this way, our approach brings a deeper level of interaction between the original and compact models.Compared to the previous knowledge distillation approaches for BERT compression, our approach does not introduce any additional loss function.Our approach outperforms existing knowledge distillation approaches on GLUE benchmark, showing a new perspective of model compression.1
Canwen Xu, Wangchunshu Zhou, Tao Ge 0001, Furu Wei, Ming Zhou 0001
EMNLP (1)3
2020 Self-Adversarial Learning with Comparative Discrimination for Text Generation
Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Furu Wei, Ming Zhou 0001
ICLR2
2020 BERT Loses Patience: Fast and Robust Inference with Early Exit
abstract
In this paper, we propose Patience-based Early Exit, a straightforward yet effective inference method that can be used as a plug-and-play technique to simultaneously improve the efficiency and robustness of a pretrained language model (PLM). To achieve this, our approach couples an internal-classifier with each layer of a PLM and dynamically stops inference when the intermediate predictions of the internal classifiers do not change for a pre-defined number of steps. Our approach improves inference efficiency as it allows the model to make a prediction with fewer layers. Meanwhile, experimental results with an ALBERT model show that our method can improve the accuracy and robustness of the model by preventing it from overthinking and exploiting multiple classifiers for prediction, yielding a better accuracy-speed trade-off compared to existing early exit methods.
Wangchunshu Zhou, Canwen Xu, Tao Ge 0001, Julian J. McAuley, Ke Xu 0001, Furu Wei
NeurIPS3
2019 Automatic Grammatical Error Correction for Sequence-to-sequence Text Generation: An Empirical Study
abstract
Sequence-to-sequence (seq2seq) models have achieved tremendous success in text generation tasks.However, there is no guarantee that they can always generate sentences without grammatical errors.In this paper, we present a preliminary empirical study on whether and how much automatic grammatical error correction can help improve seq2seq text generation.We conduct experiments across various seq2seq text generation tasks including machine translation, formality style transfer, sentence compression and simplification.Experiments show the state-of-the-art grammatical error correction system can improve the grammaticality of generated text and can bring taskoriented improvements in the tasks where target sentences are in a formal style.
Tao Ge 0001, Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001
ACL (1)1
2019 BERT-based Lexical Substitution
abstract
Previous studies on lexical substitution tend to obtain substitute candidates by finding the target word's synonyms from lexical resources (e.g., WordNet) and then rank the candidates based on its contexts.These approaches have two limitations: (1) They are likely to overlook good substitute candidates that are not the synonyms of the target words in the lexical resources;(2) They fail to take into account the substitution's influence on the global context of the sentence.To address these issues, we propose an end-toend BERT-based lexical substitution approach which can propose and validate substitute candidates without using any annotated data or manually curated resources.Our approach first applies dropout to the target word's embedding for partially masking the word, allowing BERT to take balanced consideration of the target word's semantics and contexts for proposing substitute candidates, and then validates the candidates based on their substitution's influence on the global contextualized representation of the sentence.Experiments show our approach performs well in both proposing and ranking substitute candidates, achieving the state-of-the-art results in both LS07 and LS14 benchmarks.
Wangchunshu Zhou, Tao Ge 0001, Ke Xu 0001, Furu Wei, Ming Zhou 0001
ACL (1)2
2018 Fluency Boost Learning and Inference for Neural Grammatical Error Correction
abstract
Most of the neural sequence-to-sequence (seq2seq) models for grammatical error correction (GEC) have two limitations: (1) a seq2seq model may not be well generalized with only limited error-corrected data; (2) a seq2seq model may fail to completely correct a sentence with multiple errors through normal seq2seq inference.We attempt to address these limitations by proposing a fluency boost learning and inference mechanism.Fluency boosting learning generates fluency-boost sentence pairs during training, enabling the error correction model to learn how to improve a sentence's fluency from more instances, while fluency boosting inference allows the model to correct a sentence incrementally through multi-round seq2seq inference until the sentence's fluency stops increasing.Experiments show our approaches improve the performance of seq2seq models for GEC, achieving state-of-the-art results on both CoNLL-2014 and JFLEG benchmark datasets.
Tao Ge 0001, Furu Wei, Ming Zhou 0001
ACL (1)1
2018 Fine-grained Coordinated Cross-lingual Text Stream Alignment for Endless Language Knowledge Acquisition
abstract
This paper proposes to study fine-grained coordinated cross-lingual text stream alignment through a novel information network decipherment paradigm.We use Burst Information Networks as media to represent text streams and present a simple yet effective network decipherment algorithm with diverse clues to decipher the networks for accurate text stream alignment.Experiments on Chinese-English news streams show our approach not only outperforms previous approaches on bilingual lexicon extraction from coordinated text streams but also can harvest high-quality alignments from large amounts of streaming data for endless language knowledge mining, which makes it promising to be a new paradigm for automatic language knowledge acquisition.
Tao Ge 0001, Qing Dou, Heng Ji 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001
EMNLP1
2018 EventWiki: A Knowledge Base of Major Events
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001
LREC1
2018 SeRI: A Dataset for Sub-event Relation Inference from an Encyclopedia
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Furu Wei, Ming Zhou 0001
NLPCC (2)1
2016 Event Detection with Burst Information Networks
abstract
Retrospective event detection is an important task for discovering previously unidentified events in a text stream. In this paper, we propose two fast centroid-aware event detection models based on a novel text stream representation – Burst Information Networks (BINets) for addressing the challenge. The BINets are time-aware, efficient and can be easily analyzed for identifying key information (centroids). These advantages allow the BINet-based approaches to achieve the state-of-the-art performance on multiple datasets, demonstrating the efficacy of BINets for the task of event detection.
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Zhifang Sui, Ming Zhou 0001
COLING1
2016 Towards Time-Aware Knowledge Graph Completion
abstract
Knowledge graph (KG) completion adds new facts to a KG by making inferences from existing facts. Most existing methods ignore the time information and only learn from time-unknown fact triples. In dynamic environments that evolve over time, it is important and challenging for knowledge graph completion models to take into account the temporal aspects of facts. In this paper, we present a novel time-aware knowledge graph completion model that is able to predict links in a KG using both the existing facts and the temporal information of the facts. To incorporate the happening time of facts, we propose a time-aware KG embedding model using temporal order information among facts. To incorporate the valid time of facts, we propose a joint time-aware inference model based on Integer Linear Programming (ILP) using temporal consistencyinformationasconstraints. Wefurtherintegratetwomodelstomakefulluseofglobal temporal information. We empirically evaluate our models on time-aware KG completion task. Experimental results show that our time-aware models achieve the state-of-the-art on temporal facts consistently.
Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Baobao Chang, Sujian Li, Zhifang Sui
COLING3
2016 News Stream Summarization using Burst Information Networks
abstract
This paper studies summarizing key information from news streams. We propose simple yet effective models to solve the problem based on a novel and promising representation of text streams – Burst Information Networks (BINets). A BINet can be aware of redundant information, allows global analysis of a text stream, and can be efficiently built and dynamically updated, which perfectly fits the demands of text stream summarization. Extensive experiments show that the BINet-based approaches are not only efficient and can be used in a real-time online summarization setting, but also can generate high-quality summaries, outperforming the state-of-the-art approach.
Tao Ge 0001, Lei Cui 0001, Baobao Chang, Sujian Li, Ming Zhou 0001, Zhifang Sui
EMNLP1
2016 Encoding Temporal Information for Time-Aware Link Prediction
abstract
Most existing knowledge base (KB) embedding methods solely learn from time-unknown fact triples but neglect the temporal information in the knowledge base.In this paper, we propose a novel time-aware KB embedding approach taking advantage of the happening time of facts.Specifically, we use temporal order constraints to model transformation between time-sensitive relations and enforce the embeddings to be temporally consistent and more accurate.We empirically evaluate our approach in two tasks of link prediction and triple classification.Experimental results show that our method outperforms other baselines on the two tasks consistently.
Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui
EMNLP3
2015 Bring you to the past: Automatic Generation of Topically Relevant Event Chronicles
abstract
Tao Ge, Wenzhe Pei, Heng Ji, Sujian Li, Baobao Chang, Zhifang Sui. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Tao Ge 0001, Wenzhe Pei, Heng Ji 0001, Sujian Li, Baobao Chang, Zhifang Sui
ACL (1)1
2015 An Effective Neural Network Model for Graph-based Dependency Parsing
abstract
Wenzhe Pei, Tao Ge, Baobao Chang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Wenzhe Pei, Tao Ge 0001, Baobao Chang
ACL (1)2
2015 Distinguishing Specific and Daily Topics
Tao Ge 0001, Wenzhe Pei, Baobao Chang, Zhifang Sui
APWeb1
2014 Max-Margin Tensor Neural Network for Chinese Word Segmentation
abstract
Recently, neural network models for natural language processing tasks have been increasingly focused on for their ability to alleviate the burden of manual feature engineering.In this paper, we propose a novel neural network model for Chinese word segmentation called Max-Margin Tensor Neural Network (MMTNN).By exploiting tag embeddings and tensorbased transformation, MMTNN has the ability to model complicated interactions between tags and context characters.Furthermore, a new tensor factorization approach is proposed to speed up the model and avoid overfitting.Experiments on the benchmark dataset show that our model achieves better performances than previous neural network models and that our model can achieve a competitive performance with minimal feature engineering.Despite Chinese word segmentation being a specific case, MMTNN can be easily generalized and applied to other sequence labeling tasks.
Wenzhe Pei, Tao Ge 0001, Baobao Chang
ACL (1)2
2013 Exploiting collaborative filtering techniques for automatic assessment of student free-text responses
abstract
The automatic assessment of free-text responses of students is a relatively newer task in both computational linguistics and educational technology. The goal of the task is to produce an assessment of student answers to explanation and definition questions typically asked in problems seen in practice exercises or tests. Unlike some conventional methods which assess the student responses based on only information about their corresponding questions, this paper exploits idea of collaborative filtering to analyze student responses and used an effective collaborative filtering model -- feature-based matrix factorization model to deal with this challenge. The experimental results show that our feature-based matrix factorization model outperforms the baseline models and the model with a re-ranking phase can achieve a better and competitive performance -- 63.6% overall accuracy on the Beetle dataset.
Tao Ge 0001, Zhifang Sui, Baobao Chang
CIKM1
2013 Event-Based Time Label Propagation for Automatic Dating of News Articles
abstract
Since many applications such as timeline summaries and temporal IR involving temporal analysis rely on document timestamps, the task of automatic dating of documents has been increasingly important.Instead of using feature-based methods as conventional models, our method attempts to date documents in a year level by exploiting relative temporal relations between documents and events, which are very effective for dating documents.Based on this intuition, we proposed an eventbased time label propagation model called confidence boosting in which time label information can be propagated between documents and events on a bipartite graph.The experiments show that our event-based propagation model can predict document timestamps in high accuracy and the model combined with a MaxEnt classifier outperforms the state-ofthe-art method for this task especially when the size of the training set is small.
Tao Ge 0001, Baobao Chang, Sujian Li, Zhifang Sui
EMNLP1