VLDB 2026 Research / reviewers in the wild / expert
Liying Cheng
dblp:221/0115
· DBLP profile ↗
13ranked-venue papers
4as first author
10since 2021 · last 2025
0009-0002-3071-4920ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning FrameworkabstractYew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song, Chaoqun Liu, Mahani Aljunied, Soujanya Poria, Lidong Bing |
EMNLP | 2 |
| 2024 | Exploring the Potential of Large Language Models in Computational ArgumentationabstractComputational argumentation has become an essential tool in various domains, including law, public policy, and artificial intelligence.It is an emerging research field in natural language processing that attracts increasing attention.Research on computational argumentation mainly involves two types of tasks: argument mining and argument generation.As large language models (LLMs) have demonstrated impressive capabilities in understanding context and generating natural language, it is worthwhile to evaluate the performance of LLMs on diverse computational argumentation tasks.This work aims to embark on an assessment of LLMs, such as ChatGPT, Flan models, and LLaMA2 models, in both zero-shot and few-shot settings.We organize existing tasks into six main categories and standardize the format of fourteen openly available datasets.In addition, we present a new benchmark dataset on counter speech generation that aims to holistically evaluate the end-to-end performance of LLMs on argument mining and argument generation.Extensive experiments show that LLMs exhibit commendable performance across most of the datasets, demonstrating their capabilities in the field of argumentation.Our analysis offers valuable suggestions for evaluating computational argumentation and its integration with LLMs in future research endeavors.1 Guizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong Bing |
ACL (1) | 2 |
| 2024 | Order-Agnostic Data Augmentation for Few-Shot Named Entity RecognitionabstractData augmentation (DA) methods have been proven to be effective for pre-trained language models (PLMs) in low-resource settings, including few-shot named entity recognition (NER).However, existing NER DA techniques either perform rule-based manipulations on words that break the semantic coherence of the sentence, or exploit generative models for entity or context substitution, which requires a substantial amount of labeled data and contradicts the objective of operating in low-resource settings.In this work, we propose orderagnostic data augmentation (OADA), an alternative solution that exploits the often overlooked order-agnostic property in the training data construction phase of sequence-tosequence NER methods for data augmentation.To effectively utilize the augmented data without suffering from the one-to-many issue, where multiple augmented target sequences exist for one single sentence, we further propose the use of ordering instructions and an innovative OADA-XE loss.Specifically, by treating each permutation of entity types as an ordering instruction, we rearrange the entity set accordingly, ensuring a distinct input-output pair, while OADA-XE assigns loss based on the best match between the target sequence and model predictions.We conduct comprehensive experiments and analyses across three major NER benchmarks and can significantly enhance the few-shot capabilities of PLMs with OADA.Our code is available at https://github.com/Circle-Ming/OADA-NER. Liying Cheng, Wenxuan Zhang 0001, De Wen Soh, Lidong Bing |
ACL (1) | 2 |
| 2024 | Reasoning Robustness of LLMs to Adversarial Typographical ErrorsabstractEsther Gan, Yiran Zhao, Liying Cheng, Mao Yancan, Anirudh Goyal, Kenji Kawaguchi, Min-Yen Kan, Michael Shieh. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Esther Gan, Yiran Zhao 0006, Liying Cheng, Yancan Mao, Anirudh Goyal, Kenji Kawaguchi, Min-Yen Kan, Michael Shieh |
EMNLP | 3 |
| 2024 | Large Language Models can Contrastively Refine their Generation for Better Sentence Representation LearningabstractHuiming Wang, Zhaodonghui Li, Liying Cheng, De Wen Soh, Lidong Bing. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhaodonghui Li, Liying Cheng, De Wen Soh, Lidong Bing |
NAACL-HLT | 3 |
| 2022 | IAM: A Comprehensive and Large-Scale Dataset for Integrated Argument Mining TasksabstractTraditionally, a debate usually requires a manual preparation process, including reading plenty of articles, selecting the claims, identifying the stances of the claims, seeking the evidence for the claims, etc.As the AI debate attracts more attention these years, it is worth exploring the methods to automate the tedious process involved in the debating system.In this work, we introduce a comprehensive and large dataset named IAM, which can be applied to a series of argument mining tasks, including claim extraction, stance classification, evidence extraction, etc.Our dataset is collected from over 1k articles related to 123 topics.Near 70k sentences in the dataset are fully annotated based on their argument properties (e.g., claims, stances, evidence, etc.).We further propose two new integrated argument mining tasks associated with the debate preparation process: (1) claim extraction with stance classification (CESC) and (2) claim-evidence pair extraction (CEPE).We adopt a pipeline approach and an end-to-end method for each integrated task separately.Promising experimental results are reported to show the values and challenges of our proposed tasks, and motivate future research on argument mining. 1 Liying Cheng, Lidong Bing, Ruidan He, Yan Zhang 0004, Luo Si |
ACL (1) | 1 |
| 2022 | SentBS: Sentence-level Beam Search for Controllable SummarizationabstractA wide range of control perspectives have been explored in controllable text generation.Structure-controlled summarization is recently proposed as a useful and interesting research direction.However, current structure-controlling methods have limited effectiveness in enforcing the desired structure.To address this limitation, we propose a sentence-level beam search generation method (SentBS), where evaluation is conducted throughout the generation process to select suitable sentences for subsequent generations.We experiment with different combinations of decoding methods to be used as subcomponents by SentBS and evaluate results on the structure-controlled dataset MReD.Experiments show that all explored combinations for SentBS can improve the agreement between the generated text and the desired structure, with the best method significantly reducing the structural discrepancies suffered by the existing model, by approximately 68%. 1 Chenhui Shen, Liying Cheng, Lidong Bing, Luo Si |
EMNLP | 2 |
| 2021 | Argument Pair Extraction via Attention-guided Multi-Layer Multi-Cross EncodingabstractLiying Cheng, Tianyu Wu, Lidong Bing, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Liying Cheng, Lidong Bing, Luo Si |
ACL/IJCNLP (1) | 1 |
| 2021 | On the Effectiveness of Adapter-based Tuning for Pretrained Language Model AdaptationabstractRuidan He, Linlin Liu, Hai Ye, Qingyu Tan, Bosheng Ding, Liying Cheng, Jiawei Low, Lidong Bing, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruidan He, Hai Ye, Bosheng Ding, Liying Cheng, Jia-Wei Low, Lidong Bing, Luo Si |
ACL/IJCNLP (1) | 6 |
| 2021 | Overview of Argumentative Text Understanding for AI Debater Challenge
Liying Cheng, Ruidan He, Yinzi Li, Lidong Bing, Zhongyu Wei, Qin Liu 0010, Chenhui Shen, Shuonan Zhang, Changlong Sun, Luo Si, Changjian Jiang, Xuanjing Huang 0001 |
NLPCC (2) | 2 |
| 2020 | APE: Argument Pair Extraction from Peer Review and Rebuttal via Multi-task LearningabstractPeer review and rebuttal, with rich interactions and argumentative discussions in between, are naturally a good resource to mine arguments.However, few works study both of them simultaneously.In this paper, we introduce a new argument pair extraction (APE) task on peer review and rebuttal in order to study the contents, the structure and the connections between them.We prepare a challenging dataset that contains 4,764 fully annotated review-rebuttal passage pairs from an open review platform to facilitate the study of this task.To automatically detect argumentative propositions and extract argument pairs from this corpus, we cast it as the combination of a sequence labeling task and a text relation classification task.Thus, we propose a multitask learning framework based on hierarchical LSTM networks.Extensive experiments and analysis demonstrate the effectiveness of our multi-task framework, and also show the challenges of the new task as well as motivate future research directions. 1 Liying Cheng, Lidong Bing, Wei Lu 0011, Luo Si |
EMNLP (1) | 1 |
| 2020 | ENT-DESC: Entity Description Generation by Exploring Knowledge GraphabstractPrevious works on knowledge-to-text generation take as input a few RDF triples or keyvalue pairs conveying the knowledge of some entities to generate a natural language description.Existing datasets, such as WIKIBIO, WebNLG, and E2E, basically have a good alignment between an input triple/pair set and its output text.However, in practice, the input knowledge could be more than enough, since the output description may only cover the most significant knowledge.In this paper, we introduce a large-scale and challenging dataset to facilitate the study of such a practical scenario in KG-to-text.Our dataset involves retrieving abundant knowledge of various types of main entities from a large knowledge graph (KG), which makes the current graph-to-sequence models severely suffer from the problems of information loss and parameter explosion while generating the descriptions.We address these challenges by proposing a multi-graph structure that is able to represent the original graph information more comprehensively.Furthermore, we also incorporate aggregation methods that learn to extract the rich graph information.Extensive experiments demonstrate the effectiveness of our model architecture.1 Liying Cheng, Dekun Wu, Lidong Bing, Yan Zhang 0004, Zhanming Jie, Wei Lu 0011, Luo Si |
EMNLP (1) | 1 |
| 2020 | Segmentation of pulmonary vessels based on MSFM methodabstractAccurate segmentation of pulmonary blood vessels from CT images is of great significance for lung disease detection and segmentation of other lung structures. Manual segmentation is difficult to accurately segment vascular tissue for various reasons. Therefore, in view of the existing problems and shortcomings of the existing lung vessel segmentation method, a more efficient lung vessel segmentation algorithm is proposed, that is, the multi-template fast marching method (MSFM algorithm). Firstly, it used the hole filling and maximum inter-class variance algorithm for preprocess. In the process, the lung parenchyma is extracted from the chest CT, and then the lung blood vessels are extracted in the lung parenchymal area using the MSFM algorithm. In the extraction process, the threshold and gradient are used to limit the progress of the process and the lung blood vessels are more accurately segmented. Through experimental verification, the accuracy of lung blood vessel segmentation based on MSFM algorithm is improved. Liying Cheng, Danyang Huang, Xuanshuang Gao, Daili Liang, Liuye He, Zhimei Zhang, Wenjun Tan |
HealthCom | 2 |