Chuanqi Tan

dblp:148/4497 · DBLP profile ↗
← Back
54ranked-venue papers
12as first author
31since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 10 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
abstract
Chengpeng Li, Zheng Yuan, Hongyi Yuan, Guanting Dong, Keming Lu, Jiancan Wu, Chuanqi Tan, Xiang Wang, Chang Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chengpeng Li 0001, Zheng Yuan 0002, Hongyi Yuan, Guanting Dong 0001, Keming Lu, Jiancan Wu, Chuanqi Tan, Xiang Wang 0010, Chang Zhou 0005
ACL (1)7
2024 #InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language Models
abstract
Pre-trained large language models (LLMs) can understand and align with human instructions by supervised fine-tuning (SFT). It is commonly believed that diverse and complex SFT data are of the essence to enable good instruction-following abilities. However, such diversity and complexity are obscure and lack quantitative analyses. In this work, we propose InsTag, an open-set instruction tagging method, to identify semantics and intentions of human instructions by tags that provide access to definitions and quantified analyses of instruction diversity and complexity. We obtain 6.6K fine-grained tags to describe instructions from popular open-sourced SFT datasets comprehensively. We find that the abilities of aligned LLMs benefit from more diverse and complex instructions in SFT data. Based on this observation, we propose a data sampling procedure based on InsTag, and select 6K diverse and complex samples from open-source datasets for SFT. The resulting models, TagLM, outperform open-source models based on considerably larger SFT data evaluated by MT-Bench, echoing the importance of instruction diversity and complexity and the effectiveness of InsTag. InsTag has robust potential to be extended to more applications beyond the data selection as it provides an effective way to analyze the distribution of instructions.
Keming Lu, Hongyi Yuan, Zheng Yuan 0002, Runji Lin, Junyang Lin, Chuanqi Tan, Chang Zhou 0005, Jingren Zhou 0001
ICLR6
2024 Text Diffusion Model with Encoder-Decoder Transformers for Sequence-to-Sequence Generation
abstract
Hongyi Yuan, Zheng Yuan, Chuanqi Tan, Fei Huang, Songfang Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Hongyi Yuan, Chuanqi Tan, Songfang Huang
NAACL-HLT3
2024 Sequence Labeling as Non-Autoregressive Dual-Query Set Generation
abstract
Sequence labeling is a crucial task in the NLP community that aims at identifying and assigning spans within the input sentence. It has wide applications in various fields such as information extraction, dialogue system, and sentiment analysis. However, previously proposed span-based or sequence-to-sequence models conduct locating and assigning in order, resulting in problems of error propagation and unnecessary training loss, respectively. This paper addresses the problem by reformulating the sequence labeling as a non-autoregressive set generation to realize locating and assigning in parallel. Herein, we propose aDual-QuerySetGeneration (DQSetGen) model for unified sequence labeling tasks. Specifically, the dual-query set, including a prompted type query and a positional query with anchor span, is fed into the non-autoregressive decoder to probe the spans which correspond to the positional query and have similar patterns with the type query. By avoiding the autoregressive nature of previous approaches, our method significantly improves efficiency and reduces error propagation. Experimental results illustrate that our approach can obtain superior performance on 5 sub-tasks across 11 benchmark datasets. The non-autoregressive nature of our method allows for parallel computation, achieving faster inference speed than compared baselines. In conclusion, our proposed non-autoregressive dual-query set generation method offers a more efficient and accurate approach to sequence labeling tasks in NLP. Its advantages in terms of performance and efficiency make it a promising solution for various applications in data mining and other related fields.
Xiang Chen 0016, Lei Li 0040, Shumin Deng, Chuanqi Tan, Fei Huang 0002, Luo Si, Ningyu Zhang 0001, Huajun Chen
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Reasoning with Language Model Prompting: A Survey
abstract
Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, Huajun Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shuofei Qiao, Yixin Ou, Ningyu Zhang 0001, Xiang Chen 0016, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang 0002, Huajun Chen
ACL (1)7
2023 HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation Perturbation
abstract
Language models with the Transformers structure have shown great performance in natural language processing.However, there still poses problems when fine-tuning pre-trained language models on downstream tasks, such as over-fitting or representation collapse.In this work, we propose HyPe, a simple yet effective fine-tuning technique to alleviate such problems by perturbing hidden representations of Transformers layers.Unlike previous works that only add noise to inputs or parameters, we argue that the hidden representations of Transformers layers convey more diverse and meaningful language information.Therefore, making the Transformers layers more robust to hidden representation perturbations can further benefit the fine-tuning of PLMs en bloc.We conduct extensive experiments and analyses on GLUE and other natural language inference datasets.Results demonstrate that HyPe outperforms vanilla fine-tuning and enhances generalization of hidden representations from different layers.In addition, HyPe acquires negligible computational overheads, and is better than and compatible with previous state-of-theart fine-tuning techniques.Codes are released at https://github.com/Yuanhy1997/HyPe.
Hongyi Yuan, Zheng Yuan 0002, Chuanqi Tan, Fei Huang 0002, Songfang Huang
ACL (1)3
2023 Knowledge Rumination for Pre-trained Language Models
abstract
Previous studies have revealed that vanilla pre-trained language models (PLMs) lack the capacity to handle knowledge-intensive NLP tasks alone; thus, several works have attempted to integrate external knowledge into PLMs.However, despite the promising outcome, we empirically observe that PLMs may have already encoded rich knowledge in their pre-trained parameters but fail to fully utilize them when applying them to knowledgeintensive tasks.In this paper, we propose a new paradigm dubbed Knowledge Rumination to help the pre-trained language model utilize that related latent knowledge without retrieving it from the external corpus.By simply adding a prompt like "As far as I know" to the PLMs, we try to review related latent knowledge and inject them back into the model for knowledge consolidation.We apply the proposed knowledge rumination to various language models, including RoBERTa, De-BERTa, and GPT-3.Experimental results on six commonsense reasoning tasks and GLUE benchmarks demonstrate the effectiveness of our proposed approach, which proves that the knowledge stored in PLMs can be better exploited to enhance performance 1 .
Yunzhi Yao, Peng Wang 0104, Shengyu Mao, Chuanqi Tan, Fei Huang 0002, Huajun Chen, Ningyu Zhang 0001
EMNLP4
2023 One Model for All Domains: Collaborative Domain-Prefix Tuning for Cross-Domain NER
abstract
Cross-domain NER is a challenging task to address the low-resource problem in practical scenarios. Previous typical solutions mainly obtain a NER model by pre-trained language models (PLMs) with data from a rich-resource domain and adapt it to the target domain. Owing to the mismatch issue among entity types in different domains, previous approaches normally tune all parameters of PLMs, ending up with an entirely new NER model for each domain. Moreover, current models only focus on leveraging knowledge in one general source domain while failing to successfully transfer knowledge from multiple sources to the target. To address these issues, we introduce Collaborative Domain-Prefix Tuning for cross-domain NER (CP-NER) based on text-to-text generative PLMs. Specifically, we present text-to-text generation grounding domain-related instructors to transfer knowledge to new domain NER tasks without structural modifications. We utilize frozen PLMs and conduct collaborative domain-prefix tuning to stimulate the potential of PLMs to handle NER tasks across various domains. Experimental results on the Cross-NER benchmark show that the proposed approach has flexible transfer ability and performs better on both one-source and multiple-source cross-domain NER tasks.
Xiang Chen 0016, Lei Li 0040, Shuofei Qiao, Ningyu Zhang 0001, Chuanqi Tan, Yong Jiang 0005, Fei Huang 0002, Huajun Chen
IJCAI5
2023 RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training
abstract
Vision-and-language multi-modal pretraining and fine-tuning have shown great success in visual question answering (VQA). Compared to general domain VQA, the performance of biomedical VQA suffers from limited data. In this paper, we propose a retrieval-augmented pretrain-and-finetune paradigm named RAMM for biomedical VQA to overcome the data limitation issue. Specifically, we collect a new biomedical dataset named PMCPM which offers patient-based image-text pairs containing diverse patient situations from PubMed. Then, we pretrain the biomedical multi-modal model to learn visual and textual representation for image-text pairs and align these representations with image-text contrastive objective (ITC). Finally, we propose a retrieval-augmented method to better use the limited data. We propose to retrieve similar image-text pairs based on ITC from pretraining datasets and introduce a novel retrieval-attention module to fuse the representation of the image and the question with the retrieved images and texts. Experiments demonstrate that our retrieval-augmented pretrain-and-finetune paradigm obtains state-of-the-art performance on Med-VQA2019, Med-VQA2021, VQARAD, and SLAKE datasets. Further analysis shows that the proposed RAMM and PMCPM can enhance biomedical VQA performance compared with previous resources and methods. The pre-trained models and codes are published at https://github.com/GanjinZero/RAMM.
Zheng Yuan 0005, Qiao Jin 0001, Chuanqi Tan, Zhengyun Zhao, Hongyi Yuan, Fei Huang 0002, Songfang Huang
ACM Multimedia3
2023 RRHF: Rank Responses to Align Language Models with Human Feedback
abstract
Reinforcement Learning from Human Feedback (RLHF) facilitates the alignment of large language models with human preferences, significantly enhancing the quality of interactions between humans and models. InstructGPT implements RLHF through several stages, including Supervised Fine-Tuning (SFT), reward model training, and Proximal Policy Optimization (PPO). However, PPO is sensitive to hyperparameters and requires multiple models in its standard implementation, making it hard to train and scale up to larger parameter counts. In contrast, we propose a novel learning paradigm called RRHF, which scores sampled responses from different sources via a logarithm of conditional probabilities and learns to align these probabilities with human preferences through ranking loss. RRHF can leverage sampled responses from various sources including the model responses from itself, other large language model responses, and human expert responses to learn to rank them. RRHF only needs 1 to 2 models during tuning and can efficiently align language models with human preferences robustly without complex hyperparameter tuning. Additionally, RRHF can be considered an extension of SFT and reward model training while being simpler than PPO in terms of coding, model counts, and hyperparameters. We evaluate RRHF on the Helpful and Harmless dataset, demonstrating comparable alignment performance with PPO by reward model score and human labeling. Extensive experiments show that the performance of RRHF is highly related to sampling quality which suggests RRHF is a best-of-$n$ learner.
Hongyi Yuan, Zheng Yuan 0002, Chuanqi Tan, Wei Wang 0225, Songfang Huang
NeurIPS3
2023 LOGEN: Few-Shot Logical Knowledge-Conditioned Text Generation With Self-Training
abstract
Natural language generation from structured data mainly focuses on surface-level descriptions, suffering from uncontrollable content selection and low fidelity. Previous works leverage logical forms to facilitate logical knowledge-conditioned text generation. Though achieving remarkable progress, they are data-hungry, which makes the adoption for real-world applications challenging with limited data. To this end, this paper proposes a unified framework for logical knowledge-conditioned text generation in the few-shot setting. With only a few seeds logical forms (e.g., 20/100 shot), our approach leverages self-training and samples pseudo logical forms based on content and structure consistency. Experimental results demonstrate that our approach can obtain better few-shot performance than baselines.
Shumin Deng, Hongbin Ye, Chuanqi Tan, Mosha Chen, Songfang Huang, Fei Huang 0002, Huajun Chen, Ningyu Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Learning to Ask for Data-Efficient Event Argument Extraction (Student Abstract)
abstract
Event argument extraction (EAE) is an important task for information extraction to discover specific argument roles. In this study, we cast EAE as a question-based cloze task and empirically analyze fixed discrete token template performance. As generating human-annotated question templates is often time-consuming and labor-intensive, we further propose a novel approach called “Learning to Ask,” which can learn optimized question templates for EAE without human annotations. Experiments using the ACE-2005 dataset demonstrate that our method based on optimized questions achieves state-of-the-art performance in both the few-shot and supervised settings.
Hongbin Ye, Ningyu Zhang 0001, Zhen Bi, Shumin Deng, Chuanqi Tan, Hui Chen 0018, Fei Huang 0002, Huajun Chen
AAAI5
2022 CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
abstract
Ningyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li, Xin Shang, Kangping Yin, Chuanqi Tan, Jian Xu, Fei Huang, Luo Si, Yuan Ni, Guotong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan, Linfeng Li, Jun Yan, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Ningyu Zhang 0001, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li 0040, Xin Shang, Kangping Yin, Chuanqi Tan, Fei Huang 0002, Luo Si, Yuan Ni, Guo Tong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan 0002, Jun Yan 0010, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen
ACL (1)8
2022 LightNER: A Lightweight Tuning Paradigm for Low-resource NER via Pluggable Prompting
abstract
Most NER methods rely on extensive labeled data for model training, which struggles in the low-resource scenarios with limited training data. Existing dominant approaches usually suffer from the challenge that the target domain has different label sets compared with a resource-rich source domain, which can be concluded as class transfer and domain transfer. In this paper, we propose a lightweight tuning paradigm for low-resource NER via pluggable prompting (LightNER). Specifically, we construct the unified learnable verbalizer of entity categories to generate the entity span sequence and entity categories without any label-specific classifiers, thus addressing the class transfer issue. We further propose a pluggable guidance module by incorporating learnable parameters into the self-attention layer as guidance, which can re-modulate the attention and adapt pre-trained weights. Note that we only tune those inserted module with the whole parameter of the pre-trained language model fixed, thus, making our approach lightweight and flexible for low-resource scenarios and can better transfer knowledge across domains. Experimental results show that LightNER can obtain comparable performance in the standard supervised setting and outperform strong baselines in low-resource settings.
Xiang Chen 0016, Lei Li 0040, Shumin Deng, Chuanqi Tan, Changliang Xu, Fei Huang 0002, Luo Si, Huajun Chen, Ningyu Zhang 0001
COLING4
2022 SpanProto: A Two-stage Span-based Prototypical Network for Few-shot Named Entity Recognition
abstract
Few-shot Named Entity Recognition (NER) aims to identify named entities with very little annotated data.Previous methods solve this problem based on token-wise classification, which ignores the information of entity boundaries, and inevitably the performance is affected by the massive non-entity tokens.To this end, we propose a seminal span-based prototypical network (SpanProto) that tackles few-shot NER via a two-stage approach, including span extraction and mention classification.In the span extraction stage, we transform the sequential tags into a global boundary matrix, enabling the model to focus on the explicit boundary information.For mention classification, we leverage prototypical learning to capture the semantic representations for each labeled span and make the model better adapt to novel-class entities.To further improve the model performance, we split out the false positives generated by the span extractor but not labeled in the current episode set, and then present a margin-based loss to separate them from each prototype region.Experiments over multiple benchmarks demonstrate that our model outperforms strong baselines by a large margin. 1
Jianing Wang 0002, Chengyu Wang 0001, Chuanqi Tan, Minghui Qiu, Songfang Huang, Jun Huang 0007, Ming Gao 0001
EMNLP3
2022 Differentiable Prompt Makes Pre-trained Language Models Better Few-shot Learners
Ningyu Zhang 0001, Luoqiu Li, Xiang Chen 0016, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang 0002, Huajun Chen
ICLR6
2022 Parameter-Efficient Sparsity for Large Language Models Fine-Tuning
abstract
With the dramatically increased number of parameters in language models, sparsity methods have received ever-increasing research focus to compress and accelerate the models. While most research focuses on how to accurately retain appropriate weights while maintaining the performance of the compressed model, there are challenges in the computational overhead and memory footprint of sparse training when compressing large-scale language models. To address this problem, we propose a Parameter-efficient Sparse Training (PST) method to reduce the number of trainable parameters during sparse-aware training in downstream tasks. Specifically, we first combine the data-free and data-driven criteria to efficiently and accurately measure the importance of weights. Then we investigate the intrinsic redundancy of data-driven weight importance and derive two obvious characteristics i.e. low-rankness and structuredness. Based on that, two groups of small matrices are introduced to compute the data-driven importance of weights, instead of using the original large importance score matrix, which therefore makes the sparse training resource-efficient and parameter-efficient. Experiments with diverse networks (i.e. BERT, RoBERTa and GPT-2) on dozens of datasets demonstrate PST performs on par or better than previous sparsity methods, despite only training a small number of parameters. For instance, compared with previous sparsity methods, our PST only requires 1.5% trainable parameters to achieve comparable performance on BERT.
Fuli Luo, Chuanqi Tan, Songfang Huang
IJCAI3
2022 Decoupling Knowledge from Memorization: Retrieval-augmented Prompt Learning
abstract
Prompt learning approaches have made waves in natural language processing by inducing better few-shot performance while they still follow a parametric-based learning paradigm; the oblivion and rote memorization problems in learning may encounter unstable generalization issues. Specifically, vanilla prompt learning may struggle to utilize atypical instances by rote during fully-supervised training or overfit shallow patterns with low-shot data. To alleviate such limitations, we develop RetroPrompt with the motivation of decoupling knowledge from memorization to help the model strike a balance between generalization and memorization. In contrast with vanilla prompt learning, RetroPrompt constructs an open-book knowledge-store from training instances and implements a retrieval mechanism during the process of input, training and inference, thus equipping the model with the ability to retrieve related contexts from the training corpus as cues for enhancement. Extensive experiments demonstrate that RetroPrompt can obtain better performance in both few-shot and zero-shot settings. Besides, we further illustrate that our proposed RetroPrompt can yield better generalization abilities with new datasets. Detailed analysis of memorization indeed reveals RetroPrompt can reduce the reliance of language models on memorization; thus, improving generalization for downstream tasks. Code is available in https://github.com/zjunlp/PromptKG/tree/main/research/RetroPrompt.
Xiang Chen 0016, Lei Li 0040, Ningyu Zhang 0001, Xiaozhuan Liang, Shumin Deng, Chuanqi Tan, Fei Huang 0002, Luo Si, Huajun Chen
NeurIPS6
2022 Relation Extraction as Open-book Examination: Retrieval-enhanced Prompt Tuning
abstract
Pre-trained language models have contributed significantly to relation extraction by demonstrating remarkable few-shot learning abilities. However, prompt tuning methods for relation extraction may still fail to generalize to those rare or hard patterns. Note that the previous parametric learning paradigm can be viewed as memorization regarding training data as a book and inference as the close-book test. Those long-tailed or hard patterns can hardly be memorized in parameters given few-shot instances. To this end, we regard RE as an open-book examination and propose a new semiparametric paradigm of retrieval-enhanced prompt tuning for relation extraction. We construct an open-book datastore for retrieval regarding prompt-based instance representations and corresponding relation labels as memorized key-value pairs. During inference, the model can infer relations by linearly interpolating the base output of PLM with the non-parametric nearest neighbor distribution over the datastore. In this way, our model not only infers relation through knowledge stored in the weights during training but also assists decision-making by unwinding and querying examples in the open-book datastore. Extensive experiments on benchmark datasets show that our method can achieve state-of-the-art in both standard supervised and few-shot settings
Xiang Chen 0016, Lei Li 0040, Ningyu Zhang 0001, Chuanqi Tan, Fei Huang 0002, Luo Si, Huajun Chen
SIGIR4
2022 Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph Completion
abstract
Multimodal Knowledge Graphs (MKGs), which organize visual-text factual knowledge, have recently been successfully applied to tasks such as information retrieval, question answering, and recommendation system. Since most MKGs are far from complete, extensive knowledge graph completion studies have been proposed focusing on the multimodal entity, relation extraction and link prediction. However, different tasks and modalities require changes to the model architecture, and not all images/objects are relevant to text input, which hinders the applicability to diverse real-world scenarios. In this paper, we propose a hybrid transformer with multi-level fusion to address those issues. Specifically, we leverage a hybrid transformer architecture with unified input-output for diverse multimodal knowledge graph completion tasks. Moreover, we propose multi-level fusion, which integrates visual and text representation via coarse-grained prefix-guided interaction and fine-grained correlation-aware fusion modules. We conduct extensive experiments to validate that our MKGformer can obtain SOTA performance on four datasets of multimodal link prediction, multimodal RE, and multimodal NER1. https://github.com/zjunlp/MKGformer.
Xiang Chen 0016, Ningyu Zhang 0001, Lei Li 0040, Shumin Deng, Chuanqi Tan, Changliang Xu, Fei Huang 0002, Luo Si, Huajun Chen
SIGIR5
2022 KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation Extraction
abstract
Recently, prompt-tuning has achieved promising results for specific few-shot classification tasks. The core idea of prompt-tuning is to insert text pieces (i.e., templates) into the input and transform a classification task into a masked language modeling problem. However, for relation extraction, determining an appropriate prompt template requires domain expertise, and it is cumbersome and time-consuming to obtain a suitable label word. Furthermore, there exists abundant semantic and prior knowledge among the relation labels that cannot be ignored. To this end, we focus on incorporating knowledge among relation labels into prompt-tuning for relation extraction and propose a Knowledge-aware Prompt-tuning approach with synergistic optimization (KnowPrompt). Specifically, we inject latent knowledge contained in relation labels into prompt construction with learnable virtual type words and answer words. Then, we synergistically optimize their representation with structured constraints. Extensive experimental results on five datasets with standard and low-resource settings demonstrate the effectiveness of our approach. Our code and datasets are available in GitHub1 for reproducibility.
Xiang Chen 0016, Ningyu Zhang 0001, Xin Xie 0006, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang 0002, Luo Si, Huajun Chen
WWW6
2022 Low-resource extraction with knowledge-aware pairwise prototype learning
Shumin Deng, Ningyu Zhang 0001, Hui Chen 0018, Chuanqi Tan, Fei Huang 0002, Changliang Xu, Huajun Chen
Knowl. Based Syst.4
2021 Nested Named Entity Recognition with Partially-Observed TreeCRFs
abstract
Named entity recognition (NER) is a well-studied task in natural language processing. However, the widely-used sequence labeling framework is difficult to detect entities with nested structures. In this work, we view nested NER as constituency parsing with partially-observed trees and model it with partially-observed TreeCRFs. Specifically, we view all labeled entity spans as observed nodes in a constituency tree, and other spans as latent nodes. With the TreeCRF we achieve a uniform way to jointly model the observed and the latent nodes. To compute the probability of partial trees with partial marginalization, we propose a variant of the Inside algorithm, the Masked Inside algorithm, that supports different inference operations for different nodes (evaluation for the observed, marginalization for the latent, and rejection for nodes incompatible with the observed) with efficient parallelized implementation, thus significantly speeding up training and inference. Experiments show that our approach achieves the state-of-the-art (SOTA) F1 scores on the ACE2004, ACE2005 dataset, and shows comparable performance to SOTA models on the GENIA dataset. We release the code at https://github.com/FranxYao/Partially-Observed-TreeCRFs.
Chuanqi Tan, Mosha Chen, Songfang Huang, Fei Huang 0002
AAAI2
2021 Contrastive Triple Extraction with Generative Transformer
abstract
Triple extraction is an essential task in information extraction for natural language processing and knowledge graph construction. In this paper, we revisit the end-to-end triple extraction task for sequence generation. Since generative triple extraction may struggle to capture long-term dependencies and generate unfaithful triples, we introduce a novel model, contrastive triple extraction with a generative transformer. Specifically, we introduce a single shared transformer module for encoder-decoder-based generation. To generate faithful results, we propose a novel triplet contrastive training object. Moreover, we introduce two mechanisms to further improve model performance (i.e., batch-wise dynamic attention-masking and triple-wise calibration). Experimental results on three datasets (i.e., NYT, WebNLG, and MIE) show that our approach achieves better performance than that of baselines.
Hongbin Ye, Ningyu Zhang 0001, Shumin Deng, Mosha Chen, Chuanqi Tan, Fei Huang 0002, Huajun Chen
AAAI5
2021 Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning
abstract
Recent pretrained language models extend from millions to billions of parameters.Thus the need to fine-tune an extremely large pretrained model with a limited training corpus arises in various downstream tasks.In this paper, we propose a straightforward yet effective fine-tuning technique, CHILD-TUNING, which updates a subset of parameters (called child network) of large pretrained models via strategically masking out the gradients of the non-child network during the backward process.Experiments on various downstream tasks in GLUE benchmark show that CHILD-TUNING consistently outperforms the vanilla fine-tuning by 1.5 ∼ 8.6 average score among four different pretrained models, and surpasses the prior fine-tuning techniques by 0.6 ∼ 1.3 points.Furthermore, empirical results on domain transfer and task transfer show that CHILD-TUNING can obtain better generalization performance by large margins.
Runxin Xu, Fuli Luo, Chuanqi Tan, Baobao Chang, Songfang Huang, Fei Huang 0002
EMNLP (1)4
2021 Meta Gradient Adversarial Attack
abstract
In recent years, research on adversarial attacks has be-come a hot spot. Although current literature on the transfer-based adversarial attack has achieved promising results for improving the transferability to unseen black-box models, it still leaves a long way to go. Inspired by the idea of meta-learning, this paper proposes a novel architecture called Meta Gradient Adversarial Attack (MGAA), which is plug-and-play and can be integrated with any existing gradient-based attack method for improving the cross-model transferability. Specifically, we randomly sample multiple models from a model zoo to compose different tasks and iteratively simulate a white-box attack and a black-box attack in each task. By narrowing the gap between the gradient directions in white-box and black-box attacks, the transfer-ability of adversarial examples on the black-box setting can be improved. Extensive experiments on the CIFAR10 and ImageNet datasets show that our architecture outperforms the state-of-the-art methods for both black-box and white-box attack settings.
Zheng Yuan 0005, Jie Zhang 0071, Yunpei Jia, Chuanqi Tan, Shiguang Shan
ICCV4
2021 Probing BERT in Hyperbolic Spaces
Boli Chen, Pengjun Xie, Chuanqi Tan, Mosha Chen, Liping Jing
ICLR5
2021 Document-level Relation Extraction as Semantic Segmentation
abstract
Document-level relation extraction aims to extract relations among multiple entity pairs from a document. Previously proposed graph-based or transformer-based models utilize the entities independently, regardless of global information among relational triples. This paper approaches the problem by predicting an entity-level relation matrix to capture local and global information, parallel to the semantic segmentation task in computer vision. Herein, we propose a Document U-shaped Network for document-level relation extraction. Specifically, we leverage an encoder module to capture the context information of entities and a U-shaped segmentation module over the image-style feature map to capture global interdependency among triples. Experimental results show that our approach can obtain state-of-the-art performance on three benchmark datasets DocRED, CDR, and GDA.
Ningyu Zhang 0001, Xiang Chen 0016, Xin Xie 0006, Shumin Deng, Chuanqi Tan, Mosha Chen, Fei Huang 0002, Luo Si, Huajun Chen
IJCAI5
2021 Noisy-Labeled NER with Confidence Estimation
abstract
Kun Liu, Yao Fu, Chuanqi Tan, Mosha Chen, Ningyu Zhang, Songfang Huang, Sheng Gao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chuanqi Tan, Mosha Chen, Ningyu Zhang 0001, Songfang Huang
NAACL-HLT3
2021 Contrastive Information Extraction With Generative Transformer
abstract
Information extraction tasks such as triple extraction and event extraction are of great importance for natural language processing and knowledge graph construction. In this paper, we revisit the end-to-end information extraction task for sequence generation. Since generative information extraction may struggle to capture long-term dependencies and generate unfaithful triples, we introduce a novel model, contrastive information extraction with a generative transformer. Specifically, we introduce a single shared transformer module for an encoder-decoder-based generation. To generate faithful results, we propose a novel triplet contrastive training object. Moreover, we introduce two mechanisms to further improve model performance (i.e., batch-wise dynamic attention-masking and triple-wise calibration). Experimental results on five datasets (i.e., NYT, WebNLG, MIE, ACE-2005, and MUC-4) show that our approach achieves better performance than baselines.
Ningyu Zhang 0001, Hongbin Ye, Shumin Deng, Chuanqi Tan, Mosha Chen, Songfang Huang, Fei Huang 0002, Huajun Chen
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Cross-Modal Knowledge Adaptation for Language-Based Person Search
abstract
In this paper, we present a method named Cross-Modal Knowledge Adaptation (CMKA) for language-based person search. We argue that the image and text information are not equally important in determining a person's identity. In other words, image carries image-specific information such as lighting condition and background, while text contains more modal agnostic information that is more beneficial to cross-modal matching. Based on this consideration, we propose CMKA to adapt the knowledge of image to the knowledge of text. Specially, text-to-image guidance is obtained at different levels: individuals, lists, and classes. By combining these levels of knowledge adaptation, the image-specific information is suppressed, and the common space of image and text is better constructed. We conduct experiments on the CUHK-PEDES dataset. The experimental results show that the proposed CMKA outperforms the state-of-the-art methods.
Rui Huang 0001, Hong Chang 0001, Chuanqi Tan, Bingpeng Ma
IEEE Trans. Image Process.4
2020 Boundary Enhanced Neural Span Classification for Nested Named Entity Recognition
abstract
Named entity recognition (NER) is a well-studied task in natural language processing. However, the widely-used sequence labeling framework is usually difficult to detect entities with nested structures. The span-based method that can easily detect nested entities in different subsequences is naturally suitable for the nested NER problem. However, previous span-based methods have two main issues. First, classifying all subsequences is computationally expensive and very inefficient at inference. Second, the span-based methods mainly focus on learning span representations but lack of explicit boundary supervision. To tackle the above two issues, we propose a boundary enhanced neural span classification model. In addition to classifying the span, we propose incorporating an additional boundary detection task to predict those words that are boundaries of entities. The two tasks are jointly trained under a multitask learning framework, which enhances the span representation with additional boundary supervision. In addition, the boundary detection model has the ability to generate high-quality candidate spans, which greatly reduces the time complexity during inference. Experiments show that our approach outperforms all existing methods and achieves 85.3, 83.9, and 78.3 scores in terms of F1 on the ACE2004, ACE2005, and GENIA datasets, respectively.
Chuanqi Tan, Mosha Chen, Rui Wang 0005, Fei Huang 0002
AAAI1
2020 Predicting Clinical Trial Results by Implicit Evidence Integration
abstract
Clinical trials provide essential guidance for practicing Evidence-Based Medicine, though often accompanying with unendurable costs and risks.To optimize the design of clinical trials, we introduce a novel Clinical Trial Result Prediction (CTRP) task.In the CTRP framework, a model takes a PICO-formatted clinical trial proposal with its background as input and predicts the result, i.e. how the Intervention group compares with the Comparison group in terms of the measured Outcome in the studied Population.While structured clinical evidence is prohibitively expensive for manual collection, we exploit large-scale unstructured sentences from medical literature that implicitly contain PICOs and results as evidence.Specifically, we pre-train a model to predict the disentangled results from such implicit evidence and fine-tune the model with limited data on the downstream datasets.Experiments on the benchmark Evidence Integration dataset show that the proposed model outperforms the baselines by large margins, e.g., with a 10.7% relative gain over BioBERT in macro-F1.Moreover, the performance improvement is also validated on another dataset composed of clinical trials related to COVID-19.
Qiao Jin 0001, Chuanqi Tan, Mosha Chen, Xiaozhong Liu 0001, Songfang Huang
EMNLP (1)2
2020 Latent Template Induction with Gumbel-CRFs
abstract
Learning to control the structure of sentences is a challenging problem in text generation. Existing work either relies on simple deterministic approaches or RL-based hard structures. We explore the use of structured variational autoencoders to infer latent templates for sentence generation using a soft, continuous relaxation in order to utilize reparameterization for training. Specifically, we propose a Gumbel-CRF, a continuous relaxation of the CRF sampling algorithm using a relaxed Forward-Filtering Backward-Sampling (FFBS) approach. As a reparameterized gradient estimator, the Gumbel-CRF gives more stable gradients than score-function based estimators. As a structured inference network, we show that it learns interpretable templates during training, which allows us to control the decoder during testing. We demonstrate the effectiveness of our methods with experiments on data-to-text generation and unsupervised paraphrase generation.
Chuanqi Tan, Bin Bi, Mosha Chen, Yansong Feng 0002, Alexander M. Rush
NeurIPS2
2020 Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
abstract
In this work, we propose a division-and-summarization (DaS) framework for dense video captioning. After partitioning each untrimmed long video as multiple event proposals, where each event proposal consists of a set of short video segments, we extract visual feature (e.g., C3D feature) from each segment and use the existing image/video captioning approach to generate one sentence description for this segment. Considering that the generated sentences contain rich semantic descriptions about the whole event proposal, we formulate the dense video captioning task as a visual cue aided sentence summarization problem and propose a new two stage Long Short Term Memory (LSTM) approach equipped with a new hierarchical attention mechanism to summarize all generated sentences as one descriptive sentence with the aid of visual features. Specifically, the first-stage LSTM network takes all semantic words from the generated sentences and the visual features from all segments within one event proposal as the input, and acts as the encoder to effectively summarize both semantic and visual information related to this event proposal. The second-stage LSTM network takes the output from the first-stage LSTM network and the visual features from all video segments within one event proposal as the input, and acts as the decoder to generate one descriptive sentence for this event proposal. Our comprehensive experiments on the ActivityNet Captions dataset demonstrate the effectiveness of our newly proposed DaS framework for dense video captioning.
Zhiwang Zhang, Dong Xu 0001, Wanli Ouyang, Chuanqi Tan
IEEE Trans. Circuits Syst. Video Technol.4
2019 Attention-based Transfer Learning for Brain-computer Interface
abstract
Different functional areas of the human brain play different roles in brain activity, which has not been paid sufficient research attention in the brain-computer interface (BCI) field. This paper presents a new approach for electroencephalography (EEG) classification that applies attention-based transfer learning. Our approach considers the importance of different brain functional areas to improve the accuracy of EEG classification, and provides an additional way to automatically identify brain functional areas associated with new activities without the involvement of a medical professional. We demonstrate empirically that our approach out-performs state-of-the-art approaches in the task of EEG classification, and the results of visualization indicate that our approach can detect brain functional areas related to a certain task.
Chuanqi Tan, Fuchun Sun 0001, Tao Kong, Bin Fang 0003
ICASSP1
2019 Neural Melody Composition from Lyrics
Hangbo Bao, Shaohan Huang, Furu Wei, Lei Cui 0001, Yu Wu 0012, Chuanqi Tan, Ming Zhou 0001
NLPCC (1)6
2019 A glove-based system for object recognition via visual-tactile fusion
Bin Fang 0003, Fuchun Sun 0001, Huaping Liu 0001, Chuanqi Tan, Di Guo 0002
Sci. China Inf. Sci.4
2019 LDS-FCM: A Linear Dynamical System Based Fuzzy C-Means Method for Tactile Recognition
abstract
Tactile sensing is becoming an indispensable robotic ability for object recognition and grasping manipulation despite dealing with tactile data as the force distribution over the array sensors continuously changes as a function of time. In this paper, we propose an efficient feature extractor named linear dynamic systems based fuzzy C-means method (LDS) to encode the tactile sequences, both spatially and temporally. To this end, we decompose every input sequence into multiple subsequences, each of which is locally described by a finite-ordered observability matrix of the LDS model. A fuzzy c-means method is then applied to cluster the local LDS descriptors for learning a codebook. Conditioned on the resulting codebook, the global tactile representation is formulated by employing two different frameworks to integrate the subsequences within each tactile sequence, namely, the Vector of locally aggregated descriptor and Bag-of-Word approaches. The effectiveness of the proposed model is verified by a variety of experimental evaluations on five benchmark datasets. Results reveal that our proposed method achieves a higher classification accuracy than the state-of-the-art models with a large margin.
Chunfang Liu, Wenbing Huang 0001, Fuchun Sun 0001, Minnan Luo, Chuanqi Tan
IEEE Trans. Fuzzy Syst.5
2019 Feature Pyramid Reconfiguration With Consistent Loss for Object Detection
abstract
Taking the feature pyramids into account has become a crucial way to boost the object detection performance. While various pyramid representations have been developed, previous works are still inefficient to integrate the semantical information over different scales. Moreover, recent object detectors are suffering from accurate object location applications, mainly due to the coarse definition of the "positive" examples at training and predicting phases. In this paper, we begin by analyzing current pyramid solutions, and then propose a novel architecture by reconfiguring the feature hierarchy in a flexible yet effective way. In particular, our architecture consists of two lightweight and trainable processes: global attention and local reconfiguration. The global attention is to emphasize the global information of each feature scale, while the local reconfiguration is to capture the local correlations across different scales. Both the global attention and local reconfiguration are non-linear and thus exhibit more expressive ability. Then, we discover that the loss function for object detectors during training is the central cause of the inaccurate location problem. We propose to address this issue by reshaping the standard cross entropy loss such that it focuses more on accurate predictions. Both the feature reconfiguration and the consistent loss could be utilized in popular one-stage (SSD, RetinaNet) and two-stage (Faster R-CNN) detection frameworks. Extensive experimental evaluations on PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO datasets demonstrate that, our models achieve consistent and significant boosts compared with other state-of-the-art methods.
Fuchun Sun 0001, Tao Kong, Wenbing Huang 0001, Chuanqi Tan, Bin Fang 0003, Huaping Liu 0001
IEEE Trans. Image Process.4
2018 S-Net: From Answer Extraction to Answer Synthesis for Machine Reading Comprehension
abstract
In this paper, we present a novel approach to machine reading comprehension for the MS-MARCO dataset. Unlike the SQuAD dataset that aims to answer a question with exact text spans in a passage, the MS-MARCO dataset defines the task as answering a question from multiple passages and the words in the answer are not necessary in the passages. We therefore develop an extraction-then-synthesis framework to synthesize answers from extraction results. Specifically, the answer extraction model is first employed to predict the most important sub-spans from the passage as evidence, and the answer synthesis model takes the evidence as additional features along with the question and passage to further elaborate the final answers. We build the answer extraction model with state-of-the-art neural networks for single passage reading comprehension, and propose an additional task of passage ranking to help answer extraction in multiple passages. The answer synthesis model is based on the sequence-to-sequence neural networks with extracted evidences as features. Experiments show that our extraction-then-synthesis method outperforms state-of-the-art methods.
Chuanqi Tan, Furu Wei, Nan Yang 0002, Bowen Du 0001, Weifeng Lv, Ming Zhou 0001
AAAI1
2018 A Survey on Deep Transfer Learning
Chuanqi Tan, Fuchun Sun 0001, Tao Kong, Chao Yang 0026, Chunfang Liu
ICANN (3)1
2018 Deep Transfer Learning for EEG-Based Brain Computer Interface
abstract
The electroencephalography classifier is the most important component of brain-computer interface based systems. There are two major problems hindering the improvement of it. First, traditional methods do not fully exploit multimodal information. Second, large-scale annotated EEG datasets are almost impossible to acquire because biological data acquisition is challenging and quality annotation is costly. Herein, we propose a novel deep transfer learning approach to solve these two problems. First, we model cognitive events based on EEG data by characterizing the data using EEG optical flow, which is designed to preserve multimodal EEG information in a uniform representation. Second, we design a deep transfer learning framework which is suitable for transferring knowledge by joint training, which contains a adversarial network and a special loss function. The experiments demonstrate that our approach, when applied to EEG classification tasks, has many advantages, such as robustness and accuracy.
Chuanqi Tan, Fuchun Sun 0001
ICASSP1
2018 Multiway Attention Networks for Modeling Sentence Pairs
abstract
Modeling sentence pairs plays the vital role for judging the relationship between two sentences, such as paraphrase identification, natural language inference, and answer sentence selection. Previous work achieves very promising results using neural networks with attention mechanism. In this paper, we propose the multiway attention networks which employ multiple attention functions to match sentence pairs under the matching-aggregation framework. Specifically, we design four attention functions to match words in corresponding sentences. Then, we aggregate the matching information from each function, and combine the information from all functions to obtain the final representation. Experimental results demonstrate that the proposed multiway attention networks improve the result on the Quora Question Pairs, SNLI, MultiNLI, and answer sentence selection task on the SQuAD dataset.
Chuanqi Tan, Furu Wei, Wenhui Wang 0003, Weifeng Lv, Ming Zhou 0001
IJCAI1
2018 Adaptive Adversarial Transfer Learning for Electroencephalography Classification
abstract
Insufficient training data is a serious problem in all domains related to bioinformatics. Large-scale annotated electroencephalography (EEG) datasets are almost impossible to acquire because biological data acquisition is challenging and quality annotation is costly. Transfer learning relaxes the hypothesis that the training data must be independent and identically distributed (i.i. d.) with the test data, which motivates us to use transfer learning to solve the problem of insufficient training data in bioinformatics. We propose a new approach to transfer knowledge via a deep transfer learning framework, which includes an adaptive sample selection algorithm and a joint adversarial training algorithm. The adaptive sample selection algorithm dynamically adjusts the sample weights during the training process according to the distance between the source domain and the target domain. The joint adversarial training algorithm forces the network to learn a feature extractor suitable for the target domain based on a dataset from the source domain by using an adversarial network and a specific loss function. The experiments demonstrate that our approach has many advantages, such as robustness and accuracy, when applied to EEG classification tasks.
Chuanqi Tan, Fuchun Sun 0001, Tao Kong, Chao Yang 0026, Xinyu Zhang 0001
IJCNN1
2018 Object Detection Based on Hierarchical Multi-view Proposal Network for Autonomous Driving
abstract
To achieve better results on object detection for autonomous vehicle under complex outdoor conditions, we attempt to integrated the sensor-fusion, hierarchical multi-view networks and traditional heuristical method together. The most significant environmental perception sensors for autonomous vehicles are camera and LIDAR. The 2D RGB image and 3D point cloud from camera and LIDAR respectively are utilized. The hierarchical multi-view proposal network (HMVPN) is proposed in this paper, which can effectively fuse the multi-modal information of the camera with LIDAR. As there are several hierarchical network layers in HMVPN, image becomes the input of the primary network for object detection. Moreover, LIDAR data is divided into four projection image (HBV, IBV, HCV, DCV), and then combines its original 3D point cloud into hierarchical second network to generate candidate proposals using machine learning and heuristic methods. Several simulations on the famous autonomous vehicle benchmark of KITTI show that our approach obtains about 20% higher AP than the state-of-the-art methods.
Xinyu Newman Zhang, Hongbo Gao 0001, Jialun Yin, Chuanqi Tan
IJCNN6
2018 DHA: Lidar and Vision data Fusion-based On Road Object Classifier
abstract
In this paper, we first extract three different kinds of high-level features from LIDAR point cloud, and combine them into the DHA (Depth, Height and Angle) channels. Integrated with the traditional RGB image from camera, we build a rich feature-based road object classifier by training a deep convolutional neural network model with six-channel (RGBDHA) data. Subsequently, this deep convolution neural network is fed by the integration of spacial and RGB information. With additional upsampled LIDAR data, the classifier reaches higher accuracy than single RGB image base methods. Several simulations on the famous autonomous vehicle benchmark of KITTI show that our fusion-based classifier outperforms RGB-based approaches about 15% and reaches average accuracy of 96%.
Xinyu Newman Zhang, Hongbo Gao 0001, Chuanqi Tan, Chong Xue
IJCNN5
2018 I Know There Is No Answer: Modeling Answer Validation for Machine Reading Comprehension
Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang 0002, Weifeng Lv, Ming Zhou 0001
NLPCC (1)1
2018 Context-Aware Answer Sentence Selection With Hierarchical Gated Recurrent Neural Networks
abstract
In this paper, we study the task of reading comprehension style answer sentence selection that aims to select the best sentence from a given passage to answer a question. Unlike most previous works that match the question and each candidate sentence separately, we observe that the context information among sentences in the same passage plays a vital role in this task. We propose modeling context information with hierarchical gated recurrent neural networks. Specifically, we first apply a word level recurrent neural network to model the context independent matching between the question and each candidate sentence. We then employ a sentence level recurrent neural network to incorporate the context information among all candidate sentences. Moreover, we introduce the gate mechanism to select matching information before feeding into recurrent neural networks at both word and sentence level. Experiments on the WikiQA and SQuAD datasets show that our model outperforms state-of-the-art methods.
Chuanqi Tan, Furu Wei, Qingyu Zhou, Nan Yang 0002, Bowen Du 0001, Weifeng Lv, Ming Zhou 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 Entity Linking for Queries by Searching Wikipedia Sentences
abstract
We present a simple yet effective approach for linking entities in queries.The key idea is to search sentences similar to a query from Wikipedia articles and directly use the human-annotated entities in the similar sentences as candidate entities for the query.Then, we employ a rich set of features, such as link-probability, contextmatching, word embeddings, and relatedness among candidate entities as well as their related entities, to rank the candidates under a regression based framework.The advantages of our approach lie in two aspects, which contribute to the ranking process and final linking result.First, it can greatly reduce the number of candidate entities by filtering out irrelevant entities with the words in the query.Second, we can obtain the query sensitive prior probability in addition to the static linkprobability derived from all Wikipedia articles.We conduct experiments on two benchmark datasets on entity linking for queries, namely the ERD14 dataset and the GERDAQ dataset.Experimental results show that our method outperforms state-of-the-art systems and yields 75.0% in F1 on the ERD14 dataset and 56.9% on the GERDAQ dataset.
Chuanqi Tan, Furu Wei, Pengjie Ren, Weifeng Lv, Ming Zhou 0001
EMNLP1
2017 Multimodal Classification with Deep Convolutional-Recurrent Neural Networks for Electroencephalography
Chuanqi Tan, Fuchun Sun 0001, Jianhua Chen 0009, Chunfang Liu
ICONIP (2)1
2017 Neural Question Generation from Text: A Preliminary Study
Qingyu Zhou, Nan Yang 0002, Furu Wei, Chuanqi Tan, Hangbo Bao, Ming Zhou 0001
NLPCC4
2017 A hybrid EEG-based BCI for robot grasp controlling
abstract
Brain-Computer Interfaces (BCI) can help disable people to improve human — environment interaction and rehabilitation. Grasping objects with EEG-based BCI has become a popular and hard research in recent years due to the high degree of freedom robot and complex grasp planning. Unlike commonly used paradigms, we propose a pipeline of hybrid EEG-based BCI for robot grasping by shared control to solve the key problems including target object selection, robot intelligent planning and shared control by both user intension and robot. Six experimental users could successfully use the system to grasp a number of objects in a test scene. The results of grasping experiment demonstrate that our method achieve an effective performance.
Fuchun Sun 0001, Chunfang Liu, Weihua Su, Chuanqi Tan
SMC5
2016 Solving and Generating Chinese Character Riddles
abstract
Chinese character riddle is a riddle game in which the riddle solution is a single Chinese character.It is closely connected with the shape, pronunciation or meaning of Chinese characters.The riddle description (sentence) is usually composed of phrases with rich linguistic phenomena (such as pun, simile, and metaphor), which are associated to different parts (namely radicals) of the solution character.In this paper, we propose a statistical framework to solve and generate Chinese character riddles.Specifically, we learn the alignments and rules to identify the metaphors between phrases in riddles and radicals in characters.Then, in the solving phase, we utilize a dynamic programming method to combine the identified metaphors to obtain candidate solutions.In the riddle generation phase, we use a template-based method and a replacement-based method to obtain candidate riddle descriptions.We then use Ranking SVM to rerank the candidates both in the solving and generation process.Experimental results in the solving task show that the proposed method outperforms baseline methods.We also get very promising results in the generation task according to human judges.
Chuanqi Tan, Furu Wei, Li Dong 0004, Weifeng Lv, Ming Zhou 0001
EMNLP1