Yulong Chen 0001

dblp:157/4604-1 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-2386-4656ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Benchmarking LLMs Against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
abstract
This study presents a comprehensive evaluation of the translation capabilities of existing LLMs, such as GPT-4, ALMA-R, and Deepseek-R1, compared to human translators of varying expertise levels. Through systematic human evaluation using the MQM schema, we assess translations across three language pairs (Chinese$\longleftrightarrow$English, Russian$\longleftrightarrow$English, and Chinese$\longleftrightarrow$Hindi) and three domains (News, Technology, and Biomedical). Our findings reveal that LLMs achieve performance comparable to junior-level translators in terms of total errors, while still lagging behind senior translators. Unlike traditional Neural Machine Translation systems, which show significant performance degradation in resource-poor language directions, LLMs like GPT-4 maintain consistent translation quality across all evaluated language pairs. Through qualitative analysis, we identify distinctive patterns in translation approaches: GPT-4 tends toward overly literal translations and exhibits lexical inconsistency, while human translators sometimes over-interpret context and introduce hallucinations. This study presents a systematic comparison between LLMs and human translators across different proficiency levels, providing valuable insights into the current capabilities and limitations of LLM-based translation systems.
Jianhao Yan, Pingchuan Yan, Yulong Chen 0001, Xianchao Zhu, Yue Zhang 0004
IEEE Trans. Big Data3
2025 Improving Zero-shot Sentence Decontextualisation with Content Selection and Planning
abstract
Extracting individual sentences from a document as evidence or reasoning steps is commonly done in many NLP tasks.However, extracted sentences often lack context necessary to make them understood, e.g., coreference and background information.To this end, we propose a content selection and planning framework for zero-shot decontextualisation, which determines what content should be mentioned and in what order for a sentence to be understood out of context.Specifically, given a potentially ambiguous sentence and its context, we first segment it into basic semanticallyindependent units.We then identify potentially ambiguous units from the given sentence, and extract relevant units from the context based on their discourse relations.Finally, we generate a content plan to rewrite the sentence by enriching each ambiguous unit with its relevant units.Experimental results demonstrate that our approach is competitive for sentence decontextualisation, producing sentences that exhibit better semantic integrity and discourse coherence, outperforming existing methods.
Zhenyun Deng, Yulong Chen 0001, Andreas Vlachos 0001
EMNLP2
2024 Cross-domain Constituency Parsing by Leveraging Heterogeneous Data
abstract
Knowledge transfer is investigated in various natural language processing tasks except cross-domain constituency parsing. In this paper, we leverage heterogeneous data to transfer cross-domain and cross-task knowledge to constituency parsing. Concretely, we first select language modeling, named entity recognition, CCG supertagging and dependency parsing as auxiliary tasks and collect the corpora of these tasks covering various domains as cross-domain and cross-task heterogeneous data. Second, we exploit three types of prefixes: shared, task and domain prefix, to merge cross-domain and cross-task data and decompose the general, task and domain representation in the pretrained language model. Third, we convert the data formats of multi-source heterogeneous datasets and loss objectives of the auxiliary tasks into a consistent formalization closer to constituency parsing. Finally, we jointly train the model to transfer task and domain knowledge to cross-domain constituency parsing. We verify the effectiveness of our proposed model on five target domains of MCTB. Experimental results show that our knowledge transfer model outperforms various baseline models, including conventional chart-based and transition-based parsers and the current large-scale language model for zero-shot and few-shot settings.
Peiming Guo, Meishan Zhang, Yulong Chen 0001, Jianling Li, Min Zhang 0005, Yue Zhang 0004
J. Artif. Intell. Res.3
2023 UniSumm and SummZoo: Unified Model and Diverse Benchmark for Few-Shot Summarization
abstract
Yulong Chen, Yang Liu, Ruochen Xu, Ziyi Yang, Chenguang Zhu, Michael Zeng, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yulong Chen 0001, Yang Liu 0124, Ruochen Xu, Ziyi Yang 0011, Chenguang Zhu 0001, Michael Zeng 0001, Yue Zhang 0004
ACL (1)1
2023 Revisiting Cross-Lingual Summarization: A Corpus-based Study and A New Benchmark with Improved Annotation
abstract
Yulong Chen, Huajian Zhang, Yijie Zhou, Xuefeng Bai, Yueguan Wang, Ming Zhong, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yulong Chen 0001, Xuefeng Bai 0001, Yueguan Wang, Ming Zhong 0005, Jianhao Yan, Yafu Li, Judy Li, Xianchao Zhu, Yue Zhang 0004
ACL (1)1
2023 MACSum: Controllable Summarization with Mixed Attributes
abstract
Abstract Controllable summarization allows users to generate customized summaries with specified attributes. However, due to the lack of designated annotations of controlled summaries, existing work has to craft pseudo datasets by adapting generic summarization benchmarks. Furthermore, most research focuses on controlling single attributes individually (e.g., a short summary or a highly abstractive summary) rather than controlling a mix of attributes together (e.g., a short and highly abstractive summary). In this paper, we propose MACSum, the first human-annotated summarization dataset for controlling mixed attributes. It contains source texts from two domains, news articles and dialogues, with human-annotated summaries controlled by five designed attributes (Length, Extractiveness, Specificity, Topic, and Speaker). We propose two simple and effective parameter-efficient approaches for the new task of mixed controllable summarization based on hard prompt tuning and soft prefix tuning. Results and analysis demonstrate that hard prompt models yield the best performance on most metrics and human evaluations. However, mixed-attribute control is still challenging for summarization tasks. Our dataset and code are available at https://github.com/psunlpgroup/MACSum.
Yusen Zhang 0001, Yang Liu 0124, Ziyi Yang 0011, Yuwei Fang, Yulong Chen 0001, Dragomir R. Radev, Chenguang Zhu 0001, Michael Zeng 0001, Rui Zhang 0037
Trans. Assoc. Comput. Linguistics5
2022 Graph Pre-training for AMR Parsing and Generation
abstract
meaning representation (AMR) highlights the core semantic information of text in a graph structure.Recently, pre-trained language models (PLMs) have advanced tasks of AMR parsing and AMR-to-text generation, respectively.However, PLMs are typically pretrained on textual data, thus are sub-optimal for modeling structural knowledge.To this end, we investigate graph self-supervised training to improve the structure awareness of PLMs over AMR graphs.In particular, we introduce two graph auto-encoding strategies for graphto-graph pre-training and four tasks to integrate text and graph information during pre-training.We further design a unified framework to bridge the gap between pre-training and fine-tuning tasks.Experiments on both AMR parsing and AMR-to-text generation show the superiority of our model.To our knowledge, we are the first to consider pre-training on semantic graphs.
Xuefeng Bai 0001, Yulong Chen 0001, Yue Zhang 0004
ACL (1)2
2022 Recent Advances in Text-to-SQL: A Survey of What We Have and What We Expect
abstract
Text-to-SQL has attracted attention from both the natural language processing and database communities because of its ability to convert the semantics in natural language into SQL queries and its practical application in building natural language interfaces to database systems. The major challenges in text-to-SQL lie in encoding the meaning of natural utterances, decoding to SQL queries, and translating the semantics between these two forms. These challenges have been addressed to different extents by the recent advances. However, there is still a lack of comprehensive surveys for this task. To this end, we review recent progress on text-to-SQL for datasets, methods, and evaluation and provide this systematic survey, addressing the aforementioned challenges and discussing potential future directions. We hope this survey can serve as quick access to existing work and motivate future research.
Naihao Deng, Yulong Chen 0001, Yue Zhang 0004
COLING2
2021 Semantic Representation for Dialogue Modeling
abstract
Xuefeng Bai, Yulong Chen, Linfeng Song, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xuefeng Bai 0001, Yulong Chen 0001, Linfeng Song, Yue Zhang 0004
ACL/IJCNLP (1)2
2021 On Compositional Generalization of Neural Machine Translation
abstract
Yafu Li, Yongjing Yin, Yulong Chen, Yue Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yafu Li, Yongjing Yin, Yulong Chen 0001, Yue Zhang 0004
ACL/IJCNLP (1)3
2021 DialogSum Challenge: Summarizing Real-Life Scenario Dialogues
abstract
We propose a shared task on summarizing reallife scenario dialogues, DialogSum Challenge, to encourage researchers to address challenges in dialogue summarization, which has been less studied by the summarization community.Real-life scenario dialogue summarization has a wide potential application prospect in chatbot and personal assistant.It contains unique challenges such as special discourse structure, coreference, pragmatics and social common sense, which require specific representation learning technologies to deal with.We carefully annotate a large-scale dialogue summarization dataset based on multiple public dialogue corpus, opening the door to all kinds of summarization models.
Yulong Chen 0001, Yang Liu 0124, Yue Zhang 0004
INLG1