VLDB 2026 Research / reviewers in the wild / expert
Tianyu Liu 0001
dblp:134/1099-1
· DBLP profile ↗
46ranked-venue papers
11as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 10 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | P-Aligner: Pre-Aligning LLMs via Principled Instruction Synthesis
Feifan Song 0001, Bofei Gao, Yifan Song 0002, Weimin Xiong, Yuyang Song, Tianyu Liu 0001, Houfeng Wang |
WWW | 7 |
| 2025 | Qwen2.5-xCoder: Multi-Agent Collaboration for Multilingual Code Instruction TuningabstractRecent advancement in code understanding and generation demonstrates that code LLMs fine-tuned on a high-quality instruction dataset can gain powerful capabilities to address wide-ranging code-related tasks. However, most previous existing methods mainly view each programming language in isolation and ignore the knowledge transfer among different programming languages. To bridge the gap among different programming languages, we introduce a novel multi-agent collaboration framework to enhance multilingual instruction tuning for code LLMs, where multiple language-specific intelligent agent components with generation memory work together to transfer knowledge from one language to another efficiently and effectively. Specifically, we first generate the language-specific instruction data from the code snippets and then provide the generated data as the seed data for language-specific agents. Multiple language-specific agents discuss and collaborate to formulate a new instruction and its corresponding solution (A new programming language or existing programming language), To further encourage the cross-lingual transfer, each agent stores its generation history as memory and then summarizes its merits and faults. Finally, the high-quality multilingual instruction data is used to encourage knowledge transfer among different programming languages to train Qwen2.5-xCoder. Experimental results on multilingual programming benchmarks demonstrate the superior performance of Qwen2.5-xCoder in sharing common knowledge, highlighting its potential to reduce the cross-lingual gap. Jian Yang 0003, Wei Zhang 0021, Yibo Miao, Shanghaoran Quan, Zhenhe Wu, Qiyao Peng 0006, Liqun Yang, Tianyu Liu 0001, Zeyu Cui, Binyuan Hui, Junyang Lin |
ACL (1) | 8 |
| 2025 | VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward ModelsabstractVision-language generative reward models (VL-GenRMs) play a crucial role in aligning and evaluating multimodal AI systems, yet their own evaluation remains under-explored. Current assessment methods primarily rely on AI-annotated preference labels from traditional VL tasks, which can introduce biases and often fail to effectively challenge state-of-the-art models. To address these limitations, we introduce VL-RewardBench, a comprehensive benchmark spanning general multimodal queries, visual hallucination detection, and complex reasoning tasks. Through our AI-assisted annotation pipeline that combines sample selection with human verification, we curate 1,250 high-quality examples specifically designed to probe VL-GenRMs limitations. Comprehensive evaluation across 16 leading large vision-language models demonstrates VL-RewardBench’s effectiveness as a challenging testbed, where even GPT-4o achieves only 65.4% accuracy, and state-of-the-art open-source models such as Qwen2-VL-72B, struggle to surpass random-guessing. Importantly, performance on VL-RewardBench strongly correlates (Pearson’s r > 0.9) with MMMU-Pro accuracy using Best-of-N sampling with VL-GenRMs. Analysis experiments uncover three critical insights for improving VL-GenRMs: (i) models predominantly fail at basic visual perception tasks rather than reasoning tasks; (ii) inference-time scaling benefits vary dramatically by model capacity; and (iii) training VL-GenRMs to learn to judge substantially boosts judgment capability (+14.7% accuracy for a 7B VL-GenRM). We believe VL-RewardBench along with the experimental insights will become a valuable resource for advancing VL-GenRMs. Project page: https://vl-rewardbench.github.io. Lei Li 0039, Yuancheng Wei, Zhihui Xie 0002, Xuqing Yang, Yifan Song 0002, Peiyi Wang, Chenxin An, Tianyu Liu 0001, Sujian Li, Bill Y. Lin, Lingpeng Kong, Qi Liu 0049 |
CVPR | 8 |
| 2025 | A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image GenerationabstractThis work tackles the information loss bottleneck of vector-quantization (VQ) autoregressive image generation by introducing a novel model architecture called the 2-Dimensional Autoregression (DnD) Transformer. The DnD-Transformer predicts more codes for an image by introducing a new direction, **model depth**, along with the sequence length. Compared to 1D autoregression and previous work using similar 2D image decomposition such as RQ-Transformer, the DnD-Transformer is an end-to-end model that can generate higher quality images with the same backbone model size and sequence length, opening a new optimization perspective for autoregressive image generation. Furthermore, our experiments reveal that the DnD-Transformer's potential extends beyond generating natural images. It can even generate images with rich text and graphical elements in a self-supervised manner, demonstrating an understanding of these combined modalities. This has not been previously demonstrated for popular vision generative models such as diffusion models, showing a spark of vision-language intelligence when trained solely on images. Code, datasets and models are open at https://github.com/chenllliang/DnD-Transformer. Liang Chen 0024, Sinan Tan, Zefan Cai, Weichu Xie, Haozhe Zhao, Yichi Zhang 0010, Junyang Lin, Jinze Bai, Tianyu Liu 0001, Baobao Chang |
ICLR | 9 |
| 2025 | Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsabstractRecent advancements in large language models (LLMs) have led to significant breakthroughs in mathematical reasoning capabilities.
However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To bridge this gap, we propose a comprehensive and challenging benchmark specifically designed to assess LLMs' mathematical reasoning at the Olympiad level. Unlike existing Olympiad-related benchmarks, our dataset focuses exclusively on mathematics and comprises a vast collection of 4428 competition-level problems with rigorous human annotation. These problems are meticulously categorized into over 33 sub-domains and span more than 10 distinct difficulty levels, enabling a holistic assessment of model performance in Olympiad-mathematical reasoning. Furthermore, we conducted an in-depth analysis based on this benchmark. Our experimental results show that even the most advanced models, OpenAI o1-mini and OpenAI o1-preview, struggle with highly challenging Olympiad-level problems, with 60.54% and 52.55% accuracy, highlighting significant challenges in Olympiad-level mathematical reasoning. Bofei Gao, Feifan Song 0001, Zhe Yang 0013, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li 0039, Chenghao Ma, Liang Chen 0024, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang 0009, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu 0001, Baobao Chang |
ICLR | 19 |
| 2025 | MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient EvaluationabstractJinsheng Huang, Liang Chen, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan, Haozhe Zhao, Zhihui Guo, Yichi Zhang, Jingyang Yuan, Wei Ju, Luchen Liu, Tianyu Liu, Baobao Chang, Ming Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jinsheng Huang, Liang Chen 0024, Taian Guo, Fu Zeng, Yusheng Zhao, Bohan Wu, Ye Yuan 0016, Haozhe Zhao, Zhihui Guo, Yichi Zhang 0010, Jingyang Yuan, Wei Ju 0001, Luchen Liu, Tianyu Liu 0001, Baobao Chang, Ming Zhang 0004 |
NAACL (Long Papers) | 14 |
| 2025 | CellVerse: Do Large Language Models Really Understand Cell Biology?abstractRecent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis tasks still remains unexplored. Motivated by this challenge, we introduce CellVerse, a unified language-centric question-answering benchmark that integrates four types of single-cell multi-omics data and encompasses three hierarchical levels of single-cell analysis tasks: cell type annotation (cell-level), drug response prediction (drug-level), and perturbation analysis (gene-level). Going beyond this, we systematically evaluate the performance across 14 open-source and closed-source LLMs ranging 160M $\rightarrow$ 671B on CellVerse. Remarkably, the experimental results reveal: (1) Existing specialist models (C2S-Pythia) fail to make reasonable decisions across all sub-tasks within CellVerse, while generalist models such as Qwen, Llama, GPT, and DeepSeek family models exhibit preliminary understanding capabilities within the realm of cell biology. (2) The performance of current LLMs falls short of expectations and has substantial room for improvement. Notably, in the widely studied drug response prediction task, none of the evaluated LLMs demonstrate significant performance improvement over random guessing. CellVerse offers the first large-scale empirical demonstration that significant challenges still remain in applying LLMs to cell biology. By introducing CellVerse, we lay the foundation for advancing cell biology through natural languages and hope this paradigm could facilitate next-generation single-cell analysis. Project Page: https://cellverse-cuhk.github.io Fan Zhang 0111, Tianyu Liu 0001, Zhihong Zhu 0001, Hao Wu 0094, Haixin Wang 0003, Yefeng Zheng 0001, Kun Wang 0056, Xian Wu 0001, Pheng-Ann Heng |
NeurIPS | 2 |
| 2024 | Large Language Models are not Fair EvaluatorsabstractPeiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, Zhifang Sui. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Peiyi Wang, Lei Li 0039, Liang Chen 0024, Zefan Cai, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu 0049, Tianyu Liu 0001, Zhifang Sui |
ACL (1) | 10 |
| 2024 | An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Liang Chen 0024, Haozhe Zhao, Tianyu Liu 0001, Shuai Bai, Junyang Lin, Chang Zhou 0005, Baobao Chang |
ECCV (81) | 3 |
| 2024 | Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?abstractLarge language models (LLMs) have demonstrated impressive capabilities, but still suffer from inconsistency issues (e.g.LLMs can react differently to disturbances like rephrasing or inconsequential order change).In addition to these inconsistencies, we also observe that LLMs, while capable of solving hard problems, can paradoxically fail at easier ones.To evaluate this hard-to-easy inconsistency, we develop the ConsisEval benchmark, where each entry comprises a pair of questions with a strict order of difficulty.Furthermore, we introduce the concept of consistency score to quantitatively measure this inconsistency and analyze the potential for improvement in consistency by relative consistency score.Based on comprehensive experiments across a variety of existing models, we find: (1) GPT-4 achieves the highest consistency score of 92.2% but is still inconsistent to specific questions due to distraction by redundant information, misinterpretation of questions, etc.; (2) models with stronger capabilities typically exhibit higher consistency, but exceptions also exist; (3) hard data enhances consistency for both fine-tuning and in-context learning.Our data and code will be publicly available on GitHub. 1 Zhe Yang 0013, Yichang Zhang, Tianyu Liu 0001, Jian Yang 0003, Junyang Lin, Chang Zhou 0005, Zhifang Sui |
EMNLP | 3 |
| 2024 | DialogVCS: Robust Natural Language Understanding in Dialogue System UpgradeabstractZefan Cai, Xin Zheng, Tianyu Liu, Haoran Meng, Jiaqi Han, Gang Yuan, Binghuai Lin, Baobao Chang, Yunbo Cao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zefan Cai, Tianyu Liu 0001, Haoran Meng, Binghuai Lin, Baobao Chang, Yunbo Cao |
NAACL-HLT | 3 |
| 2023 | Denoising Bottleneck with Mutual Information Maximization for Video Multimodal FusionabstractShaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu, Binghuai Lin, Yunbo Cao, Zhifang Sui. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui |
ACL (1) | 4 |
| 2022 | Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual CluesabstractQingxiu Dong, Ziwei Qin, Heming Xia, Tian Feng, Shoujie Tong, Haoran Meng, Lin Xu, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu, Zhifang Sui. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Qingxiu Dong, Ziwei Qin, Heming Xia, Shoujie Tong, Haoran Meng, Zhongyu Wei, Weidong Zhan, Baobao Chang, Sujian Li, Tianyu Liu 0001, Zhifang Sui |
ACL (1) | 12 |
| 2022 | A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationabstractTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, Bill Dolan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Tianyu Liu 0001, Yizhe Zhang 0002, Chris Brockett, Zhifang Sui, Weizhu Chen, William B. Dolan |
ACL (1) | 1 |
| 2022 | Learning Robust Representations for Continual Relation Extraction via Adversarial Class AugmentationabstractContinual relation extraction (CRE) aims to continually learn new relations from a classincremental data stream.CRE model usually suffers from catastrophic forgetting problem, i.e., the performance of old relations seriously degrades when the model learns new relations.Most previous work attributes catastrophic forgetting to the corruption of the learned representations as new relations come, with an implicit assumption that the CRE models have adequately learned the old relations.In this paper, through empirical studies we argue that this assumption may not hold, and an important reason for catastrophic forgetting is that the learned representations do not have good robustness against the appearance of analogous relations in the subsequent learning process.To address this issue, we encourage the model to learn more precise and robust representations through a simple yet effective adversarial class augmentation mechanism (ACA), which is easy to implement and model-agnostic.Experimental results show that ACA can consistently improve the performance of state-of-theart CRE models on two popular benchmarks. Peiyi Wang, Yifan Song 0002, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Sujian Li, Zhifang Sui |
EMNLP | 3 |
| 2022 | HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex label hierarchy.Recently, the pretrained language models (PLM) have been widely adopted in HTC through a finetuning paradigm.However, in this paradigm, there exists a huge gap between the classification tasks with sophisticated label hierarchy and the masked language model (MLM) pretraining tasks of PLMs and thus the potential of PLMs cannot be fully tapped.To bridge the gap, in this paper, we propose HPT, a Hierarchy-aware Prompt Tuning method to handle HTC from a multi-label MLM perspective.Specifically, we construct a dynamic virtual template and label words that take the form of soft prompts to fuse the label hierarchy knowledge and introduce a zero-bounded multi-label cross-entropy loss to harmonize the objectives of HTC and MLM.Extensive experiments show HPT achieves state-of-the-art performances on 3 popular HTC datasets and is adept at handling the imbalance and low resource situations. Peiyi Wang, Tianyu Liu 0001, Binghuai Lin, Yunbo Cao, Zhifang Sui, Houfeng Wang |
EMNLP | 3 |
| 2022 | Robust Fine-tuning via Perturbation and Interpolation from In-batch InstancesabstractFine-tuning pretrained language models (PLMs) on downstream tasks has become common practice in natural language processing. However, most of the PLMs are vulnerable, e.g., they are brittle under adversarial attacks or imbalanced data, which hinders the application of the PLMs on some downstream tasks, especially in safe-critical scenarios. In this paper, we propose a simple yet effective fine-tuning method called Match-Tuning to force the PLMs to be more robust. For each instance in a batch, we involve other instances in the same batch to interact with it. To be specific, regarding the instances with other labels as a perturbation, Match-Tuning makes the model more robust to noise at the beginning of training. While nearing the end, Match-Tuning focuses more on performing an interpolation among the instances with the same label for better generalization. Extensive experiments on various tasks in GLUE benchmark show that Match-Tuning consistently outperforms the vanilla fine-tuning by 1.64 scores. Moreover, Match-Tuning exhibits remarkable robustness to adversarial attacks and data imbalance. Shoujie Tong, Qingxiu Dong, Damai Dai, Yifan Song 0002, Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
IJCAI | 5 |
| 2022 | An Enhanced Span-based Decomposition Method for Few-Shot Sequence LabelingabstractPeiyi Wang, Runxin Xu, Tianyu Liu, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Qingyu Zhou, Yunbo Cao, Baobao Chang, Zhifang Sui |
NAACL-HLT | 3 |
| 2022 | A Two-Stream AMR-enhanced Model for Document-level Event Argument ExtractionabstractRunxin Xu, Peiyi Wang, Tianyu Liu, Shuang Zeng, Baobao Chang, Zhifang Sui. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Runxin Xu, Peiyi Wang, Tianyu Liu 0001, Shuang Zeng, Baobao Chang, Zhifang Sui |
NAACL-HLT | 3 |
| 2021 | Towards Faithfulness in Open Domain Table-to-text Generation from an Entity-centric ViewabstractIn open domain table-to-text generation, we notice the unfaithful generation usually contains hallucinated entities which can not be aligned to any input table record. We thus try to evaluate the generation faithfulness with two entity-centric metrics: table record coverage and the ratio of hallucinated entities in text, both of which are shown to have strong agreement with human judgements. Then based on these metrics, we quantitatively analyze the correlation between training data quality and generation fidelity which indicates the potential usage of entity information in faithful generation. Motivated by these findings, we propose two methods for faithful generation: 1) augmented training by incorporating the auxiliary entity information, including both an augmented plan-based model and an unsupervised model and 2) training instance selection based on faithfulness ranking. We show these approaches improve generation fidelity in both full dataset setting and few shot setting by both automatic and human evaluations. Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
AAAI | 1 |
| 2021 | Document-level Event Extraction via Heterogeneous Graph-based Interaction Model with a TrackerabstractRunxin Xu, Tianyu Liu, Lei Li, Baobao Chang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Runxin Xu, Tianyu Liu 0001, Lei Li 0005, Baobao Chang |
ACL/IJCNLP (1) | 2 |
| 2021 | Behind the Scenes: An Exploration of Trigger Biases Problem in Few-Shot Event ClassificationabstractFew-Shot Event Classification (FSEC) aims at developing a model for event prediction, which can generalize to new event types with a limited number of annotated data. Existing FSEC studies have achieved high accuracy on different benchmarks. However, we find they suffer from trigger biases that signify the statistical homogeneity between some trigger words and target event types, which we summarize as trigger overlapping and trigger separability. The biases can result in context-bypassing problem, i.e., correct classifications can be gained by looking at only the trigger words while ignoring the entire context. Therefore, existing models can be weak in generalizing to unseen data in real scenarios. To further uncover the trigger biases and assess the generalization ability of the models, we propose two new sampling methods, Trigger-Uniform Sampling (TUS) and COnfusion Sampling (COS), for the meta tasks construction during evaluation. Besides, to cope with the context-bypassing problem in FSEC models, we introduce adversarial training and trigger reconstruction techniques. Experiments show these techniques help not only improve the performance, but also enhance the generalization ability of models. Peiyi Wang, Runxin Xu, Tianyu Liu 0001, Damai Dai, Baobao Chang, Zhifang Sui |
CIKM | 3 |
| 2021 | A Privacy-Preserving Incentive Mechanism for Federated Cloud-Edge LearningabstractThe federated learning scheme enhances the privacy preservation through avoiding the private data uploading in cloud-edge computing. However, the attacks against the uploaded model updates still cause private data leakage which demotivates the privacy-sensitive participating edge devices. Facing this issue, we aim to design a privacy-preserving incentive mechanism for the federated cloud-edge learning (PFCEL) system such that 1) the edge devices are motivated to actively contribute to the updated model uploading, 2) a trade-off between the private data leakage and the model accuracy is achieved. We formulate the incentive design problem as a three-layer Stackelberg game, where the server-device interaction is further formulated as a contract design problem. Extensive numerical evaluations demonstrate the effectiveness of our designed mechanism in terms of privacy preservation and system utility. Tianyu Liu 0001, Boya Di, Lingyang Song |
GLOBECOM | 1 |
| 2021 | Decompose, Fuse and Generate: A Formation-Informed Method for Chinese Definition GenerationabstractHua Zheng, Damai Dai, Lei Li, Tianyu Liu, Zhifang Sui, Baobao Chang, Yang Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Damai Dai, Lei Li 0039, Tianyu Liu 0001, Zhifang Sui, Baobao Chang, Yang Liu 0124 |
NAACL-HLT | 4 |
| 2020 | An Anchor-Based Automatic Evaluation Metric for Document SummarizationabstractThe widespread adoption of reference-based automatic evaluation metrics such as ROUGE has promoted the development of document summarization.In this paper, we consider a new protocol for designing reference-based metrics that require the endorsement of source document(s).Following protocol, we propose an anchored ROUGE metric fixing each summary particle on source document, which bases the computation on more solid ground.Empirical results on benchmark datasets validate that source document helps to induce a higher correlation with human judgments for ROUGE metric.Being self-explanatory and easy-to-implement, the protocol can naturally foster various effective designs of reference-based metrics besides the anchored ROUGE introduced here. Kexiang Wang, Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
COLING | 2 |
| 2020 | An Empirical Study on Model-agnostic Debiasing Strategies for Robust Natural Language InferenceabstractThe prior work on natural language inference (NLI) debiasing mainly targets at one or few known biases while not necessarily making the models more robust.In this paper, we focus on the model-agnostic debiasing strategies and explore how to (or is it possible to) make the NLI models robust to multiple distinct adversarial attacks while keeping or even strengthening the models' generalization power.We firstly benchmark prevailing neural NLI models including pretrained ones on various adversarial datasets.We then try to combat distinct known biases by modifying a mixture of experts (MoE) ensemble method (Clark et al., 2019) and show that it's nontrivial to mitigate multiple NLI biases at the same time, and that model-level ensemble method outperforms MoE ensemble method.We also perform data augmentation including text swap, word substitution and paraphrase and prove its efficiency in combating various (though not all) adversarial attacks at the same time.Finally, we investigate several methods to merge heterogeneous training data (1.35M) and perform model ensembling, which are straightforward but effective to strengthen NLI models. Tianyu Liu 0001, Xiaoan Ding, Baobao Chang, Zhifang Sui |
CoNLL | 1 |
| 2020 | Discriminatively-Tuned Generative Classifiers for Robust Natural Language InferenceabstractWhile discriminative neural network classifiers are generally preferred, recent work has shown advantages of generative classifiers in term of data efficiency and robustness.In this paper, we focus on natural language inference (NLI).We propose GenNLI, a generative classifier for NLI tasks, and empirically characterize its performance by comparing it to five baselines, including discriminative models and large-scale pretrained language representation models like BERT.We explore training objectives for discriminative fine-tuning of our generative classifiers, showing improvements over log loss fine-tuning from prior work (Lewis and Fan, 2019).In particular, we find strong results with a simple unbounded modification to log loss, which we call the "infinilog loss".Our experiments show that GenNLI outperforms both discriminative and pretrained baselines across several challenging NLI experimental settings, including small training sets, imbalanced label distributions, and label noise. Xiaoan Ding, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Kevin Gimpel |
EMNLP (1) | 2 |
| 2020 | An Exploration of Arbitrary-Order Sequence Labeling via Energy-Based Inference NetworksabstractMany tasks in natural language processing involve predicting structured outputs, e.g., sequence labeling, semantic role labeling, parsing, and machine translation.Researchers are increasingly applying deep representation learning to these problems, but the structured component of these approaches is usually quite simplistic.In this work, we propose several high-order energy terms to capture complex dependencies among labels in sequence labeling, including several that consider the entire label sequence.We use neural parameterizations for these energy terms, drawing from convolutional, recurrent, and selfattention networks.We use the framework of learning energy-based inference networks (Tu and Gimpel, 2018) for dealing with the difficulties of training and inference with such models.We empirically demonstrate that this approach achieves substantial improvement using a variety of high-order energy terms on four sequence labeling tasks, while having the same decoding speed as simple, local classifiers.We also find high-order energies to help in noisy data conditions. 1 Lifu Tu, Tianyu Liu 0001, Kevin Gimpel |
EMNLP (1) | 2 |
| 2020 | HypoNLI: Exploring the Artificial Patterns of Hypothesis-only Bias in Natural Language InferenceabstractMany recent studies have shown that for models trained on datasets for natural language inference (NLI), it is possible to make correct predictions by merely looking at the hypothesis while completely ignoring the premise. In this work, we manage to derive adversarial examples in terms of the hypothesis-only bias and explore eligible ways to mitigate such bias. Specifically, we extract various phrases from the hypotheses (artificial patterns) in the training sets, and show that they have been strong indicators to the specific labels. We then figure out ‘hard’ and ‘easy’ instances from the original test sets whose labels are opposite to or consistent with those indications. We also set up baselines including both pretrained models (BERT, RoBerta, XLNet) and competitive non-pretrained models (InferSent, DAM, ESIM). Apart from the benchmark and baselines, we also investigate two debiasing approaches which exploit the artificial pattern modeling to mitigate such hypothesis-only bias: down-sampling and adversarial training. We believe those methods can be treated as competitive baselines in NLI debiasing tasks. Tianyu Liu 0001, Baobao Chang, Zhifang Sui |
LREC | 1 |
| 2019 | Hierarchical Encoder with Auxiliary Supervision for Neural Table-to-Text Generation: Learning Better Representation for TablesabstractGenerating natural language descriptions for the structured tables which consist of multiple attribute-value tuples is a convenient way to help people to understand the tables. Most neural table-to-text models are based on the encoder-decoder framework. However, it is hard for a vanilla encoder to learn the accurate semantic representation of a complex table. The challenges are two-fold: firstly, the table-to-text datasets often contain large number of attributes across different domains, thus it is hard for the encoder to incorporate these heterogeneous resources. Secondly, the single encoder also has difficulties in modeling the complex attribute-value structure of the tables. To this end, we first propose a two-level hierarchical encoder with coarse-to-fine attention to handle the attribute-value structure of the tables. Furthermore, to capture the accurate semantic representations of the tables, we propose 3 joint tasks apart from the prime encoder-decoder learning, namely auxiliary sequence labeling task, text autoencoder and multi-labeling classification, as the auxiliary supervisions for the table encoder. We test our models on the widely used dataset WIKIBIO which contains Wikipedia infoboxes and related descriptions. The dataset contains complex tables as well as large number of attributes across different domains. We achieve the state-of-the-art performance on both automatic and human evaluation metrics. Tianyu Liu 0001, Fuli Luo, Qiaolin Xia, Shuming Ma, Baobao Chang, Zhifang Sui |
AAAI | 1 |
| 2019 | Towards Comprehensive Description Generation from Factual Attribute-value TablesabstractThe comprehensive descriptions for factual attribute-value tables, which should be accurate, informative and loyal, can be very helpful for end users to understand the structured data in this form.However previous neural generators might suffer from key attributes missing, less informative and groundless information problems, which impede the generation of high-quality comprehensive descriptions for tables.To relieve these problems, we first propose force attention (FA) method to encourage the generator to pay more attention to the uncovered attributes to avoid potential key attributes missing.Furthermore, we propose reinforcement learning for information richness to generate more informative as well as more loyal descriptions for tables.In our experiments, we utilize the widely used WIKIBIO dataset as a benchmark.Additionally we create WB-filter based on WIKIBIO to test our model in the simulated user-oriented scenarios, in which the generated descriptions should accord with particular user interests.Experimental results show that our model outperforms the state-of-the-art baselines on both automatic and human evaluation. Tianyu Liu 0001, Fuli Luo, Wei Wu 0044, Baobao Chang, Zhifang Sui |
ACL (1) | 1 |
| 2019 | Learning to Control the Fine-grained Sentiment for Story Ending GenerationabstractAutomatic story ending generation is an interesting and challenging task in natural language generation.Previous studies are mainly limited to generate coherent, reasonable and diversified story endings, and few works focus on controlling the sentiment of story endings.This paper focuses on generating a story ending which meets the given fine-grained sentiment intensity.There are two major challenges to this task.First is the lack of story corpus which has fine-grained sentiment labels.Second is the difficulty of explicitly controlling sentiment intensity when generating endings.Therefore, we propose a generic and novel framework which consists of a sentiment analyzer and a sentimental generator, respectively addressing the two challenges.The sentiment analyzer adopts a series of methods to acquire sentiment intensities of the story dataset.The sentimental generator introduces the sentiment intensity into decoder via a Gaussian Kernel Layer to control the sentiment of the output.To the best of our knowledge, this is the first endeavor to control the fine-grained sentiment for story ending generation without manually annotating sentiment labels.Experiments show that our proposed framework can generate story endings which are not only more coherent and fluent but also able to meet the given sentiment intensity better. 1 Fuli Luo, Damai Dai, Tianyu Liu 0001, Baobao Chang, Zhifang Sui, Xu Sun 0001 |
ACL (1) | 4 |
| 2019 | Key Fact as Pivot: A Two-Stage Model for Low Resource Table-to-Text GenerationabstractTable -to-text generation aims to translate the structured data into the unstructured text.Most existing methods adopt the encoder-decoder framework to learn the transformation, which requires large-scale training samples.However, the lack of large parallel data is a major practical problem for many domains.In this work, we consider the scenario of low resource table-to-text generation, where only limited parallel data is available.We propose a novel model to separate the generation into two stages: key fact prediction and surface realization.It first predicts the key facts from the tables, and then generates the text with the key facts.The training of key fact prediction needs much fewer annotated data, while surface realization can be trained with pseudo parallel corpus.We evaluate our model on a biography generation dataset.Our model can achieve 27.34 BLEU score with only 1, 000 parallel data, while the baseline model only obtain the performance of 9.71 BLEU score. 1 Shuming Ma, Tianyu Liu 0001, Peng Li 0030, Jie Zhou 0016, Xu Sun 0001 |
ACL (1) | 3 |
| 2019 | MAAM: A Morphology-Aware Alignment Model for Unsupervised Bilingual Lexicon InductionabstractThe task of unsupervised bilingual lexicon induction (UBLI) aims to induce word translations from monolingual corpora in two languages.Previous work has shown that morphological variation is an intractable challenge for the UBLI task, where the induced translation in failure case is usually morphologically related to the correct translation.To tackle this challenge, we propose a morphology-aware alignment model for the UBLI task.The proposed model aims to alleviate the adverse effect of morphological variation by introducing grammatical information learned by the pre-trained denoising language model.Results show that our approach can substantially outperform several state-of-the-art unsupervised systems, and even achieves competitive performance compared to supervised methods. Fuli Luo, Tianyu Liu 0001, Xu Sun 0001 |
ACL (1) | 4 |
| 2019 | Enhancing Topic-to-Essay Generation with External Commonsense KnowledgeabstractAutomatic topic-to-essay generation is a challenging task since it requires generating novel, diverse, and topic-consistent paragraph-level text with a set of topics as input.Previous work tends to perform essay generation based solely on the given topics while ignoring massive commonsense knowledge.However, this commonsense knowledge provides additional background information, which can help to generate essays that are more novel and diverse.Towards filling this gap, we propose to integrate commonsense from the external knowledge base into the generator through dynamic memory mechanism.Besides, the adversarial training based on a multi-label discriminator is employed to further improve topic-consistency.We also develop a series of automatic evaluation metrics to comprehensively assess the quality of the generated essay.Experiments show that with external commonsense knowledge and adversarial training, the generated essays are more novel, diverse, and topic-consistent than existing methods in terms of both automatic and human evaluation. Lei Li 0039, Fuli Luo, Tianyu Liu 0001, Xu Sun 0001 |
ACL (1) | 4 |
| 2019 | Playing Card-Based RTS Games with Deep Reinforcement LearningabstractGame AI is of great importance as games are simulations of reality. Recent research on game AI has shown much progress in various kinds of games, such as console games, board games and MOBA games. However, the exploration in RTS games remains a challenge for their huge state space, imperfect information, sparse rewards and various strategies. Besides, the typical card-based RTS games have complex card features and are still lacking solutions. We present a deep model SEAT (selection-attention) to play card-based RTS games. The SEAT model includes two parts, a selection part for card choice and an attention part for card usage, and it learns from scratch via deep reinforcement learning. Comprehensive experiments are performed on Clash Royale, a popular mobile card-based RTS game. Empirical results show that the SEAT model agent makes it to reach a high winning rate against rule-based agents and decision-tree-based agent. Tianyu Liu 0001, Hongchang Li, Kaigui Bian, Lingyang Song |
IJCAI | 1 |
| 2018 | Table-to-Text Generation by Structure-Aware Seq2seq LearningabstractTable-to-text generation aims to generate a description for a factual table which can be viewed as a set of field-value records. To encode both the content and the structure of a table, we propose a novel structure-aware seq2seq architecture which consists of field-gating encoder and description generator with dual attention. In the encoding phase, we update the cell memory of the LSTM unit by a field gate and its corresponding field value in order to incorporate field information into table representation. In the decoding phase, dual attention mechanism which contains word level attention and field level attention is proposed to model the semantic relevance between the generated description and the table. We conduct experiments on the WIKIBIO dataset which contains over 700k biographies and corresponding infoboxes from Wikipedia. The attention visualizations and case studies show that our model is capable of generating coherent and informative descriptions based on the comprehensive understanding of both the content and the structure of a table. Automatic evaluations also show our model outperforms the baselines by a great margin. Code for this work is available on https://github.com/tyliupku/wiki2bio. Tianyu Liu 0001, Kexiang Wang, Lei Sha, Baobao Chang, Zhifang Sui |
AAAI | 1 |
| 2018 | Order-Planning Neural Text Generation From Structured DataabstractGenerating texts from structured data (e.g., a table) is important for various natural language processing tasks such as question answering and dialog systems. In recent studies, researchers use neural language models and encoder-decoder frameworks for table-to-text generation. However, these neural network-based approaches typically do not model the order of content during text generation. When a human writes a summary based on a given table, he or she would probably consider the content order before wording. In this paper, we propose an order-planning text generation model, where order information is explicitly captured by link-based attention. Then a self-adaptive gate combines the link-based attention with traditional content-based attention. We conducted experiments on the WikiBio dataset and achieve higher performance than previous methods in terms of BLEU, ROUGE, and NIST scores; we also performed ablation tests to analyze each component of our model. Lei Sha, Lili Mou, Tianyu Liu 0001, Pascal Poupart, Sujian Li, Baobao Chang, Zhifang Sui |
AAAI | 3 |
| 2018 | Incorporating Glosses into Neural Word Sense DisambiguationabstractWord Sense Disambiguation (WSD) aims to identify the correct meaning of polysemous words in the particular context.Lexical resources like WordNet which are proved to be of great help for WSD in the knowledge-based methods.However, previous neural networks for WSD always rely on massive labeled data (context), ignoring lexical resources like glosses (sense definitions).In this paper, we integrate the context and glosses of the target word into a unified framework in order to make full use of both labeled data and lexical knowledge.Therefore, we propose GAS: a gloss-augmented WSD neural network which jointly encodes the context and glosses of the target word.GAS models the semantic relationship between the context and the gloss in an improved memory network framework, which breaks the barriers of the previous supervised methods and knowledge-based methods.We further extend the original gloss of word sense via its semantic relations in WordNet to enrich the gloss information.The experimental results show that our model outperforms the state-of-theart systems on several English all-words WSD datasets. Fuli Luo, Tianyu Liu 0001, Qiaolin Xia, Baobao Chang, Zhifang Sui |
ACL (1) | 2 |
| 2018 | Leveraging Gloss Knowledge in Neural Word Sense Disambiguation by Hierarchical Co-AttentionabstractThe goal of Word Sense Disambiguation (WSD) is to identify the correct meaning of a word in the particular context.Traditional supervised methods only use labeled data (context), while missing rich lexical knowledge such as the gloss which defines the meaning of a word sense.Recent studies have shown that incorporating glosses into neural networks for WSD has made significant improvement.However, the previous models usually build the context representation and gloss representation separately.In this paper, we find that the learning for the context and gloss representation can benefit from each other.Gloss can help to highlight the important words in the context, thus building a better context representation.Context can also help to locate the key words in the gloss of the correct word sense.Therefore, we introduce a co-attention mechanism to generate co-dependent representations for the context and gloss.Furthermore, in order to capture both word-level and sentence-level information, we extend the attention mechanism in a hierarchical fashion.Experimental results show that our model achieves the state-of-the-art results on several standard English all-words WSD test datasets. Fuli Luo, Tianyu Liu 0001, Zexue He, Qiaolin Xia, Zhifang Sui, Baobao Chang |
EMNLP | 2 |
| 2018 | Phrase-level Self-Attention Networks for Universal Sentence EncodingabstractUniversal sentence encoding is a hot topic in recent NLP research.Attention mechanism has been an integral part in many sentence encoding models, allowing the models to capture context dependencies regardless of the distance between elements in the sequence.Fully attention-based models have recently attracted enormous interest due to their highly parallelizable computation and significantly less training time.However, the memory consumption of their models grows quadratically with sentence length, and the syntactic information is neglected.To this end, we propose Phrase-level Self-Attention Networks (PSAN) that perform self-attention across words inside a phrase to capture context dependencies at the phrase level, and use the gated memory updating mechanism to refine each word's representation hierarchically with longer-term context dependencies captured in a larger phrase.As a result, the memory consumption can be reduced because the self-attention is performed at the phrase level instead of the sentence level.At the same time, syntactic information can be easily integrated in the model.Experiment results show that PSAN can achieve the state-ofthe-art transfer performance across a plethora of NLP tasks including sentence classification, natural language inference and sentence textual similarity. Wei Wu 0044, Houfeng Wang, Tianyu Liu 0001, Shuming Ma |
EMNLP | 3 |
| 2017 | A Soft-label Method for Noise-tolerant Distantly Supervised Relation ExtractionabstractDistant-supervised relation extraction inevitably suffers from wrong labeling problems because it heuristically labels relational facts with knowledge bases.Previous sentence level denoise models don't achieve satisfying performances because they use hard labels which are determined by distant supervision and immutable during training.To this end, we introduce an entity-pair level denoise method which exploits semantic information from correctly labeled entity pairs to correct wrong labels dynamically during training.We propose a joint score function which combines the relational scores based on the entity-pair representation and the confidence of the hard label to obtain a new label, namely a soft label, for certain entity pair.During training, soft labels instead of hard labels serve as gold labels.Experiments on the benchmark dataset show that our method dramatically reduces noisy instances and outperforms the state-of-the-art systems. Tianyu Liu 0001, Kexiang Wang, Baobao Chang, Zhifang Sui |
EMNLP | 1 |
| 2017 | Affinity-Preserving Random Walk for Multi-Document SummarizationabstractMulti-document summarization provides users with a short text that summarizes the information in a set of related documents.This paper introduces affinitypreserving random walk to the summarization task, which preserves the affinity relations of sentences by an absorbing random walk model.Meanwhile, we put forward adjustable affinity-preserving random walk to enforce the diversity constraint of summarization in the random walk process.The ROUGE evaluations on DUC 2003 topic-focused summarization task and DUC 2004 generic summarization task show the good performance of our method, which has the best ROUGE-2 recall among the graph-based ranking methods. Kexiang Wang, Tianyu Liu 0001, Zhifang Sui, Baobao Chang |
EMNLP | 2 |
| 2017 | Large-Scale Simple Question Generation by Template-Based Seq2seq Learning
Tianyu Liu 0001, Bingzhen Wei, Baobao Chang, Zhifang Sui |
NLPCC | 1 |
| 2016 | Towards Time-Aware Knowledge Graph CompletionabstractKnowledge graph (KG) completion adds new facts to a KG by making inferences from existing facts. Most existing methods ignore the time information and only learn from time-unknown fact triples. In dynamic environments that evolve over time, it is important and challenging for knowledge graph completion models to take into account the temporal aspects of facts. In this paper, we present a novel time-aware knowledge graph completion model that is able to predict links in a KG using both the existing facts and the temporal information of the facts. To incorporate the happening time of facts, we propose a time-aware KG embedding model using temporal order information among facts. To incorporate the valid time of facts, we propose a joint time-aware inference model based on Integer Linear Programming (ILP) using temporal consistencyinformationasconstraints. Wefurtherintegratetwomodelstomakefulluseofglobal temporal information. We empirically evaluate our models on time-aware KG completion task. Experimental results show that our time-aware models achieve the state-of-the-art on temporal facts consistently. Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Baobao Chang, Sujian Li, Zhifang Sui |
COLING | 2 |
| 2016 | Encoding Temporal Information for Time-Aware Link PredictionabstractMost existing knowledge base (KB) embedding methods solely learn from time-unknown fact triples but neglect the temporal information in the knowledge base.In this paper, we propose a novel time-aware KB embedding approach taking advantage of the happening time of facts.Specifically, we use temporal order constraints to model transformation between time-sensitive relations and enforce the embeddings to be temporally consistent and more accurate.We empirically evaluate our approach in two tasks of link prediction and triple classification.Experimental results show that our method outperforms other baselines on the two tasks consistently. Tingsong Jiang, Tianyu Liu 0001, Tao Ge 0001, Lei Sha, Sujian Li, Baobao Chang, Zhifang Sui |
EMNLP | 2 |